Method and device for analyzing system network threat risk and electronic equipment
By breaking down network threat risks into multi-dimensional threat factors and performing weighted correlation analysis, a network threat reasoning report is generated. This addresses the shortcomings of the passive analysis framework in existing technologies, enabling proactive identification and evolution tracking of complex network threats, and improving the accuracy and real-time performance of security situation awareness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ELECTRONICS CORP 6TH RES INST
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies, when dealing with complex network threats, suffer from passive rule matching mechanisms that result in a lack of ability to discover unknown threats. Single-dimensional analysis frameworks lack proactive correlation capabilities, and traditional threat intelligence lacks proactive update mechanisms, making it difficult to proactively discover new threats and form a complete proactive discovery loop, causing threat analysis to lag behind attack evolution.
By decomposing the network threat risks of the system to be analyzed into multi-dimensional threat factors, a dataset of clue elements including clue elements and their weights is generated. Weighted correlation analysis is then performed to mine target risk clue element groups and generate network threat reasoning reports, enabling proactive identification and evolution tracking of complex network threats.
It improves the accuracy and real-time performance of the system's network security situational awareness, effectively responds to new threats such as advanced persistent threats and supply chain attacks, forms a proactive discovery loop, and meets the proactive defense needs of large-scale network threats.
Smart Images

Figure CN121907579A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular to a method, apparatus, and electronic device for analyzing system network threat risks. Background Technology
[0002] With the rapid development of information technology, the methods of cyberattacks against systems are becoming increasingly complex. In particular, new threats such as advanced persistent threats (APS) and supply chain attacks pose a severe challenge to traditional security protection systems. Traditional security protection tools (such as firewalls and intrusion detection systems) are no longer able to cope with new threats such as APS, supply chain attacks, and vulnerability exploits. Existing methods for analyzing network threat risks usually adopt a passive threat analysis framework and rely on known vulnerability databases, log anomaly detection, or attack chain modeling. Moreover, risk correlation analysis algorithms are generally isolated from applications, making it impossible to actively explore leads from initial clues and detect hidden threat paths, resulting in threat analysis lagging behind the speed of attack evolution.
[0003] Existing technologies suffer from several shortcomings when dealing with complex network threats, including: passive rule-matching mechanisms leading to a lack of unknown threat detection capabilities and an inability to proactively uncover new threats; a lack of proactive correlation capabilities in single-dimensional analysis frameworks; software analysis failing to proactively differentiate the potential risks of self-developed or outsourced code; vendor analysis failing to proactively correlate integrator technology stacks and third-party library dependencies; isolated algorithm applications lacking proactive collaboration mechanisms, resulting in low efficiency and a lack of proactive outreach capabilities when processing large-scale data; and traditional threat intelligence lacking proactive update mechanisms, making it difficult to support the dynamic improvement of correlation and outreach. In summary, existing technologies lack the ability to proactively explore and extend threat networks from initial clues, making it difficult to form a complete proactive discovery loop. This causes threat analysis to lag behind attack evolution, failing to meet the needs of proactive defense against large-scale network threats. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method, apparatus, and electronic device for analyzing system network threat risks. By decomposing the network threat risks existing in the system to be analyzed into multi-dimensional threat factors, a dataset of clue elements including clue elements and their weights is generated. Weighted correlation analysis is then performed on the clue element dataset to mine target risk clue element groups. Based on the target risk clue element groups, a network threat reasoning report corresponding to the system to be analyzed is generated, realizing the proactive identification and evolution tracking of complex network threats. While deeply analyzing the components of the system and their interrelationships, it effectively responds to new threats such as advanced persistent threats and supply chain attacks, improving the accuracy and real-time performance of system network security situation awareness.
[0005] This application provides a method for analyzing system network threat risks, the method comprising: Obtain system information corresponding to the system to be analyzed, and based on the system information, decompose and weight the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed; Based on the set of threat factors and the correlation between each threat factor in the set of threat factors, the dataset of clue elements corresponding to the system to be analyzed is determined. A weighted association analysis is performed on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset; Based on the target risk clue element group, a network threat reasoning report corresponding to the system to be analyzed is generated to analyze the network threat risk of the system to be analyzed.
[0006] Furthermore, the step of obtaining system information corresponding to the system to be analyzed, and based on the system information, splitting and weighting the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed, includes: Obtain system information corresponding to the system to be analyzed; wherein, the system information includes at least hardware component information, software information, supplier information, and developer behavior information; Based on the system information, the network threats of the system to be analyzed are split into threat factors, and the hardware layer threat factors, software layer threat factors, supply chain layer threat factors and developer layer threat factors corresponding to the system to be analyzed are determined respectively. Based on the hardware score value corresponding to the hardware component information and multiple preset hardware factor coefficients, the hardware layer factor weight value corresponding to the hardware layer threat factor is determined. Based on the software score corresponding to the software information and multiple preset software factor coefficients, determine the software layer factor weight value corresponding to the software layer threat factor; Based on the supplier rating value corresponding to the supplier information and multiple preset supply chain factor coefficients, the supply chain layer factor weight value corresponding to the supply chain layer threat factor is determined. Based on the developer rating value corresponding to the developer behavior information and multiple preset developer factor coefficients, the developer layer factor weight value corresponding to the developer layer threat factor is determined. The hardware layer threat factor, the hardware layer factor weight value, the software layer threat factor, the software layer factor weight value, the supply chain layer threat factor, the supply chain layer factor weight value, the developer layer threat factor, and the developer layer factor weight value are determined as the set of threat factors corresponding to the system to be analyzed.
[0007] Furthermore, determining the dataset of clue elements corresponding to the system to be analyzed based on the set of threat factors and the correlation between each threat factor in the set of threat factors includes: Obtain the correlation between each threat factor in the threat factor set; Based on the aforementioned relationships, data mining is performed on the set of threat factors to identify multiple clue elements composed of the threat factors, and a preset graph database tool is used to determine the association and combination relationships between the clue elements. For each of the aforementioned clue elements, a clue weight value corresponding to that clue element is determined based on the factor weight value corresponding to each of the aforementioned threat factors. The clue elements, the associated combination relationships, and the clue weight value corresponding to each clue element are determined as the clue element dataset corresponding to the system to be analyzed.
[0008] Furthermore, the step of performing weighted association analysis on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset includes: The dataset of clue elements is scanned to extract a first frequent itemset from the dataset, and the clue weight value and association combination relationship corresponding to each clue element in the first frequent itemset are determined. The clue elements in the first frequent itemset are sorted in descending order according to the numerical value of the clue weight to obtain the first frequent itemset sequence; Based on the pre-defined root node, the first frequent item sequence, the cue weight value, and the association combination relationship between the cue elements in the first frequent item set, a frequent pattern tree corresponding to the system to be analyzed is constructed by recursion. Frequent pattern growth mining is performed on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed.
[0009] Furthermore, the step of performing frequent pattern growth mining on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed includes: Determine whether there is only one target path in the frequent pattern tree; If there is only one target path in the frequent pattern tree, then the nodes corresponding to all clue elements in the frequent pattern tree are added to the target path, so that the union of the frequent pattern tree and the first frequent itemset is determined as the target frequent itemset. If there are multiple target paths in the frequent pattern tree, then traverse the nodes corresponding to all clue elements in the frequent pattern tree to extract multiple conditional pattern bases related to the first frequent itemset, and calculate the weighted support corresponding to each conditional pattern base. Each weighted support is compared with a preset support threshold, and the conditional pattern bases corresponding to the weighted support that are greater than or equal to the support threshold are selected to obtain the target frequent itemset composed of the conditional pattern bases. Calculate the target weighted support for each clue element in the target frequent itemset, and the weighted confidence for each group of clue elements in the target frequent itemset. Based on the target weighted support and the weighted confidence, determine the target score value corresponding to each clue element group; Based on the target score, at least one group of target risk clue elements corresponding to the system to be analyzed is determined from the target frequent itemset.
[0010] Furthermore, the step of generating a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group includes: Threat and risk correlation information is extracted from the target risk clue element group using a pre-defined knowledge graph. Based on the threat risk association information, the attack path information and risk information corresponding to the system to be analyzed are determined respectively; Based on the attack path information and the risk information, mitigation suggestions are determined for the system to be analyzed. Based on the attack path information, the risk information, and the mitigation suggestion information, a network threat inference report corresponding to the system to be analyzed is generated.
[0011] Furthermore, the step of generating a network threat inference report corresponding to the system to be analyzed based on the target risk clue element group also includes: In response to a user obtaining external intelligence data, the knowledge graph is updated and replaced based on the external intelligence data, the network threat reasoning report, and the target risk clue element group to obtain an updated knowledge graph.
[0012] Furthermore, after generating the network threat inference report corresponding to the system to be analyzed, the analysis method further includes: In response to the system under analysis triggering at least one preset clue lifecycle management condition, the system under analysis performs the operation corresponding to the triggered clue lifecycle management condition; wherein, the clue lifecycle management condition includes clue generation condition, association activation condition, weight update condition, clue merging or downgrading condition, and knowledge accumulation condition.
[0013] This application embodiment also provides a system network threat risk analysis device, the analysis device comprising: The threat factor splitting module is used to obtain system information corresponding to the system to be analyzed, and based on the system information, to split and weight the network threats of the system to be analyzed, and determine the set of threat factors corresponding to the system to be analyzed. The clue element generation module is used to determine the clue element dataset corresponding to the system to be analyzed based on the threat factor set and the correlation between each threat factor in the threat factor set; The weighted correlation analysis module is used to perform weighted correlation analysis on the clue element dataset and determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset. The inference report generation module is used to generate a network threat inference report corresponding to the system to be analyzed based on the target risk clue element group, so as to analyze the network threat risk of the system to be analyzed.
[0014] Furthermore, when the threat factor splitting module is used to obtain system information corresponding to the system to be analyzed, and based on the system information, to split and weight the network threats of the system to be analyzed, and determine the set of threat factors corresponding to the system to be analyzed, the threat factor splitting module is used to: Obtain system information corresponding to the system to be analyzed; wherein, the system information includes at least hardware component information, software information, supplier information, and developer behavior information; Based on the system information, the network threats of the system to be analyzed are split into threat factors, and the hardware layer threat factors, software layer threat factors, supply chain layer threat factors and developer layer threat factors corresponding to the system to be analyzed are determined respectively. Based on the hardware score value corresponding to the hardware component information and multiple preset hardware factor coefficients, the hardware layer factor weight value corresponding to the hardware layer threat factor is determined. Based on the software score corresponding to the software information and multiple preset software factor coefficients, determine the software layer factor weight value corresponding to the software layer threat factor; Based on the supplier rating value corresponding to the supplier information and multiple preset supply chain factor coefficients, the supply chain layer factor weight value corresponding to the supply chain layer threat factor is determined. Based on the developer rating value corresponding to the developer behavior information and multiple preset developer factor coefficients, the developer layer factor weight value corresponding to the developer layer threat factor is determined. The hardware layer threat factor, the hardware layer factor weight value, the software layer threat factor, the software layer factor weight value, the supply chain layer threat factor, the supply chain layer factor weight value, the developer layer threat factor, and the developer layer factor weight value are determined as the set of threat factors corresponding to the system to be analyzed.
[0015] Furthermore, when the clue element generation module determines the clue element dataset corresponding to the system to be analyzed based on the threat factor set and the correlation between each threat factor in the threat factor set, the clue element generation module is used to: Obtain the correlation between each threat factor in the threat factor set; Based on the aforementioned relationships, data mining is performed on the set of threat factors to identify multiple clue elements composed of the threat factors, and a preset graph database tool is used to determine the association and combination relationships between the clue elements. For each of the aforementioned clue elements, a clue weight value corresponding to that clue element is determined based on the factor weight value corresponding to each of the aforementioned threat factors. The clue elements, the associated combination relationships, and the clue weight value corresponding to each clue element are determined as the clue element dataset corresponding to the system to be analyzed.
[0016] Furthermore, when the weighted correlation analysis module performs weighted correlation analysis on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset, the weighted correlation analysis module is used to: The dataset of clue elements is scanned to extract a first frequent itemset from the dataset, and the clue weight value and association combination relationship corresponding to each clue element in the first frequent itemset are determined. The clue elements in the first frequent itemset are sorted in descending order according to the numerical value of the clue weight to obtain the first frequent itemset sequence; Based on the pre-defined root node, the first frequent item sequence, the cue weight value, and the association combination relationship between the cue elements in the first frequent item set, a frequent pattern tree corresponding to the system to be analyzed is constructed by recursion. Frequent pattern growth mining is performed on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed.
[0017] Furthermore, when the weighted association analysis module is used to perform frequent pattern growth mining on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed, the weighted association analysis module is used for: Determine whether there is only one target path in the frequent pattern tree; If there is only one target path in the frequent pattern tree, then the nodes corresponding to all clue elements in the frequent pattern tree are added to the target path, so that the union of the frequent pattern tree and the first frequent itemset is determined as the target frequent itemset. If there are multiple target paths in the frequent pattern tree, then traverse the nodes corresponding to all clue elements in the frequent pattern tree to extract multiple conditional pattern bases related to the first frequent itemset, and calculate the weighted support corresponding to each conditional pattern base. Each weighted support is compared with a preset support threshold, and the conditional pattern bases corresponding to the weighted support that are greater than or equal to the support threshold are selected to obtain the target frequent itemset composed of the conditional pattern bases. Calculate the target weighted support for each clue element in the target frequent itemset, and the weighted confidence for each group of clue elements in the target frequent itemset. Based on the target weighted support and the weighted confidence, determine the target score value corresponding to each clue element group; Based on the target score, at least one group of target risk clue elements corresponding to the system to be analyzed is determined from the target frequent itemset.
[0018] Furthermore, when the reasoning report generation module generates a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group, the reasoning report generation module is used to: Threat and risk correlation information is extracted from the target risk clue element group using a pre-defined knowledge graph. Based on the threat risk association information, the attack path information and risk information corresponding to the system to be analyzed are determined respectively; Based on the attack path information and the risk information, mitigation suggestions are determined for the system to be analyzed. Based on the attack path information, the risk information, and the mitigation suggestion information, a network threat inference report corresponding to the system to be analyzed is generated.
[0019] Furthermore, when the reasoning report generation module generates a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group, the reasoning report generation module is also used for: In response to a user obtaining external intelligence data, the knowledge graph is updated and replaced based on the external intelligence data, the network threat reasoning report, and the target risk clue element group to obtain an updated knowledge graph.
[0020] Furthermore, the analysis device also includes a lifecycle management module, which is used for: In response to the system under analysis triggering at least one preset clue lifecycle management condition, the system under analysis performs the operation corresponding to the triggered clue lifecycle management condition; wherein, the clue lifecycle management condition includes clue generation condition, association activation condition, weight update condition, clue merging or downgrading condition, and knowledge accumulation condition.
[0021] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the system network threat risk analysis method described above are performed.
[0022] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the system network threat risk analysis method described above.
[0023] The present application provides a method, apparatus, and electronic device for analyzing network threat risks in a system. The analysis method includes: acquiring system information corresponding to a system to be analyzed; and based on the system information, splitting and weighting the network threats of the system to be analyzed to determine a set of threat factors corresponding to the system to be analyzed; determining a dataset of clue elements corresponding to the system to be analyzed based on the set of threat factors and the correlation between each threat factor in the set of threat factors; performing weighted correlation analysis on the dataset of clue elements to determine at least one group of target risk clue elements corresponding to the system to be analyzed; and generating a network threat reasoning report corresponding to the system to be analyzed based on the group of target risk clue elements to analyze the network threat risks of the system to be analyzed.
[0024] Compared to existing technologies that employ passive threat analysis frameworks and rely on known vulnerability databases, log anomaly detection, or attack chain modeling, and whose risk correlation analysis algorithms are generally isolated, this approach decomposes the network threat risks of the system under analysis into multi-dimensional threat factors, generates a dataset of clue elements including clue elements and their weights, performs weighted correlation analysis on the clue element dataset to uncover target risk clue element groups, and then generates a network threat inference report corresponding to the system under analysis based on the target risk clue element groups. This enables proactive identification and evolution tracking of complex network threats, effectively responds to new threats such as advanced persistent threats and supply chain attacks while deeply analyzing the system's components and their interrelationships, and improves the accuracy and real-time performance of the system's network security situational awareness.
[0025] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 One of the flowcharts for a system network threat risk analysis method provided in this application embodiment; Figure 2 A second flowchart illustrating a method for analyzing system network threat risks provided in this application embodiment; Figure 3 One of the structural schematic diagrams of a system network threat risk analysis device provided in an embodiment of this application; Figure 4 A second schematic diagram of a system network threat risk analysis device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.
[0029] Research has revealed that with the rapid development of information technology, the methods of cyberattacks against systems are becoming increasingly complex. In particular, emerging threats such as advanced persistent threats (APS) and supply chain attacks pose a severe challenge to traditional security protection systems. Traditional security protection tools (such as firewalls and intrusion detection systems) are no longer sufficient to cope with emerging threats such as APS, supply chain attacks, and vulnerability exploits. Existing methods for analyzing network threat risks typically employ a passive threat analysis framework and rely on known vulnerability databases, log anomaly detection, or attack chain modeling. Furthermore, risk correlation analysis algorithms are generally isolated from applications, making it impossible to proactively explore leads from initial clues and detect hidden threat paths, resulting in threat analysis lagging behind the speed of attack evolution.
[0030] Existing technologies suffer from several shortcomings when dealing with complex network threats, including: limitations of static analysis (traditional tools rely on known vulnerability databases for matching, failing to proactively identify undisclosed threat paths, especially difficult to discover hidden related threats within the supply chain); lack of dynamic evolution capabilities (existing systems are mostly based on static asset lists, making it difficult to adapt to the dynamic characteristics of threat factors changing over time and in the environment. For example, when a supplier is exposed for a security incident or a vulnerability is publicly disclosed, traditional systems cannot automatically update the relevant threat weights and reassess the risk); lack of a weighting mechanism (traditional association rule mining does not consider the differentiated weights of threat factors, leading to the neglect of high-risk, low-frequency combinations); weak multi-dimensional analysis capabilities (traditional tools only focus on the network or code layer, failing to integrate multi-dimensional factors such as hardware, supply chain, and development teams); and insufficient knowledge accumulation and reuse (existing technologies lack effective mechanisms for accumulating and reusing threat analysis results, failing to form a continuously evolving knowledge system, resulting in the need for re-analysis when similar threats reappear, reducing analysis efficiency).
[0031] Based on the aforementioned issues, existing technologies have the following shortcomings in addressing complex network threats: passive rule-matching mechanisms result in a lack of ability to discover unknown threats and are unable to proactively uncover new threats; single-dimensional analysis frameworks lack proactive correlation capabilities; software analysis fails to proactively distinguish the potential risks of self-developed or outsourced code; vendor analysis fails to proactively correlate integrator technology stacks and third-party library dependencies; isolated algorithm applications lack proactive collaboration mechanisms, resulting in low efficiency and a lack of proactive outreach capabilities when processing large-scale data; traditional threat intelligence lacks proactive update mechanisms, making it difficult to support the dynamic improvement of correlation and outreach. In summary, existing technologies lack the ability to proactively explore and extend leads from initial clue elements, making it difficult to form a complete proactive discovery loop, causing threat analysis to lag behind attack evolution and failing to meet the needs of proactive defense against large-scale network threats.
[0032] Based on this, this application provides a method for analyzing network threat risks in a system. By decomposing the network threat risks existing in the system to be analyzed into multi-dimensional threat factors, a dataset of clue elements including clue elements and their weights is generated. Weighted correlation analysis is then performed on the dataset of clue elements to mine target risk clue element groups. Based on the target risk clue element groups, a network threat reasoning report corresponding to the system to be analyzed is generated, enabling proactive identification and evolution tracking of complex network threats. While deeply analyzing the components of the system and their interrelationships, this method effectively addresses new threats such as advanced persistent threats and supply chain attacks, improving the accuracy and real-time performance of system network security situation awareness.
[0033] Please see Figure 1 , Figure 1 This is one of the flowcharts for a system network threat risk analysis method provided in an embodiment of this application. For example... Figure 1 As shown in the embodiments of this application, the method for analyzing system network threat risks includes: S101. Obtain system information corresponding to the system to be analyzed, and based on the system information, split and weight the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed.
[0034] It should be noted that the system to be analyzed refers to a system that is expected to undergo network threat risk analysis using the method described in the embodiments of this application.
[0035] Here, the system to be analyzed may include a critical information infrastructure (CII) system. For example, the system to be analyzed includes, but is not limited to, network-connected systems and equipment such as critical infrastructure, financial systems, and enterprise information systems.
[0036] In this embodiment of the application, the purpose of threat factor decomposition is to break down the complex system to be analyzed into multiple quantifiable threat factors, which facilitates subsequent correlation analysis and risk assessment. A multi-level and multi-dimensional decomposition method is adopted to decompose the system to be analyzed from four main levels: hardware, software, supply chain and developers, to form a complete set of threat factors.
[0037] The system information includes at least hardware component information, software information, supplier information, and developer behavior information.
[0038] In this embodiment of the application, the set of threat factors includes hardware layer threat factors, hardware layer factor weight values corresponding to hardware layer threat factors, software layer threat factors, software layer factor weight values corresponding to software layer threat factors, supply chain layer threat factors, supply chain layer factor weight values corresponding to supply chain layer threat factors, developer layer threat factors, and developer layer factor weight values corresponding to developer layer threat factors.
[0039] In one possible implementation of this application, step S101 may include: S1011. Obtain the system information corresponding to the system to be analyzed.
[0040] In this embodiment of the application, the hardware component information refers to the hardware component information of the computer device in the system to be analyzed, including but not limited to CPU model, network card MAC address, OUI manufacturer information, etc.
[0041] The software information includes, but is not limited to, self-developed and outsourced code, technology stack, and third-party library versions.
[0042] The supplier information refers to the supplier's basic information, including but not limited to the supplier's size, business scope, and geographical location.
[0043] The developer behavior information includes, but is not limited to, behavioral data such as the developer's historical project risk ratings and code submission records, as well as the collection of the developer's identity information, namely, the developer's role, permissions, work experience, etc.
[0044] S1012. Based on the system information, the network threats of the system to be analyzed are split into threat factors, and the hardware layer threat factors, software layer threat factors, supply chain layer threat factors and developer layer threat factors corresponding to the system to be analyzed are determined respectively.
[0045] S1013. Based on the hardware score value corresponding to the hardware component information and the preset multiple hardware factor coefficients, determine the hardware layer factor weight value corresponding to the hardware layer threat factor.
[0046] In this step, risk scores are assigned to the OUI vendor information of hardware components. The security risk level of a vendor is assessed by querying the vendor's security history database (such as NIST CSF 2.0 compliance records, historical security incidents, etc.). If a vendor's hardware devices have been exposed to supply chain attack vulnerabilities, its risk score will be increased accordingly.
[0047] Here, the hardware score corresponding to the hardware component information includes the hardware component type risk score, the manufacturer's historical security incident score, and the hardware configuration compliance score.
[0048] Among them, the risk score for hardware component type is determined based on the criticality of the component in the system (such as CPU, network card, storage device, etc.); the score for vendor historical security incidents is determined by querying publicly available security incident databases (such as CISA alerts, CVE vulnerability databases) to count the number of historical security issues of the vendor; and the score for hardware configuration compliance is assessed by the degree of matching with security standards.
[0049] In this embodiment of the application, the hardware layer factor weight value corresponding to the hardware layer threat factor is calculated using the following formula.
[0050] .
[0051] in, This represents the hardware layer factor weight value; , , These represent multiple preset hardware factor coefficients (e.g., , , (These can be set to 0.3, 0.2, or 0.5 respectively). This indicates the risk score for the type of hardware component; This indicates the manufacturer's historical security incident score; This indicates the hardware configuration compliance score.
[0052] S1014. Based on the software score value corresponding to the software information and the preset multiple software factor coefficients, determine the software layer factor weight value corresponding to the software layer threat factor.
[0053] In this step, code scanning tools are used to analyze the code repository and identify the ratio of self-developed code to outsourced code; software technology stack information, including programming languages, frameworks, libraries, etc., is extracted and matched with known high-risk technology stack databases; version information of third-party libraries is analyzed to check for known vulnerabilities (such as CVE vulnerability databases).
[0054] Here, the software score corresponding to the software information includes the number of third-party library vulnerabilities, the proportion of open-source components, and the risk level of the technology stack.
[0055] The number of vulnerabilities in third-party libraries is obtained by synchronizing the CVE vulnerability database in real time; the proportion of open-source components is determined by using code scanning tools to count the ratio of open-source code to self-developed code; and the risk level of the technology stack is assessed based on the degree of exposure of the technology stack in historical attacks.
[0056] In this embodiment of the application, the software layer factor weight value corresponding to the software layer threat factor is calculated using the following formula.
[0057] .
[0058] in, This represents the software layer factor weight value; Indicates the number of vulnerabilities in third-party libraries; Indicates the percentage of open-source components; Indicates the risk level of the technology stack; , , These represent multiple preset software factor coefficients (e.g., , , These can be set to 0.4, 0.3, and 0.3 respectively.
[0059] S1015. Based on the supplier rating value corresponding to the supplier information and the preset multiple supply chain factor coefficients, determine the supply chain layer factor weight value corresponding to the supply chain layer threat factor.
[0060] In this step, based on supplier information, the supplier's compliance is assessed, including whether it meets the preset security standards; the supplier's financial status and operational risks are analyzed, and data is obtained through publicly available financial reports and industry analysis; the supplier's historical security incidents are assessed, including past security vulnerabilities and data breaches.
[0061] Here, the supplier rating corresponding to the supplier information includes policy compliance rating, financial risk rating, and operational risk rating.
[0062] In this embodiment of the application, the supply chain layer factor weight value corresponding to the supply chain layer threat factor is calculated by the following formula.
[0063] .
[0064] in, This represents the weight values of factors at the supply chain level; This indicates the policy compliance score; This indicates the financial risk score; This indicates the operational risk score; , , These represent multiple preset supply chain factor coefficients (e.g., , , These can be set to 0.4, 0.3, and 0.3 respectively.
[0065] S1016. Based on the developer rating value corresponding to the developer behavior information and the preset multiple developer factor coefficients, determine the developer layer factor weight value corresponding to the developer layer threat factor.
[0066] In this step, based on developer behavior information, the risks of the developer's historical projects are assessed by analyzing the number and severity of security issues that occurred in the projects the developer participated in; the developer's code commit records are analyzed, including commit frequency, commit time, and commit content, to identify abnormal behavior (such as a large number of commits outside of working hours, frequent modifications to critical code, etc.).
[0067] Here, the developer rating corresponding to the developer behavior information includes historical project risk rating and code submission anomaly level.
[0068] Among them, the risk rating of historical projects is determined by analyzing the vulnerability patching history and security audit results of developers involved in the projects; the code submission anomaly is calculated by statistically analyzing the time, frequency, content and other characteristics of the submission records; and the security awareness score is evaluated based on indicators such as the developer's participation in security training and initiative in reporting vulnerabilities.
[0069] In this embodiment of the application, the developer layer factor weight value corresponding to the developer layer threat factor is calculated by the following formula.
[0070] .
[0071] in, This represents the developer-level factor weight value; Indicates the risk rating of historical projects; Indicates the degree of code submission exception; and Each represents a coefficient of a developer factor (e.g., and These can be set to 0.6 and 0.4 respectively.
[0072] S1017. The hardware layer threat factor, the hardware layer factor weight value, the software layer threat factor, the software layer factor weight value, the supply chain layer threat factor, the supply chain layer factor weight value, the developer layer threat factor, and the developer layer factor weight value are determined as the set of threat factors corresponding to the system to be analyzed.
[0073] In this step, after the threat factors are decomposed, the threat factors and their weight values are stored in the clue database (SPANDB) to form a set of threat factors.
[0074] The threat factor set includes the threat type, associated entities, and factor weight values for each threat factor, providing a foundation for subsequent correlation analysis and risk assessment.
[0075] S102. Based on the set of threat factors and the correlation between each threat factor in the set of threat factors, determine the dataset of clue elements corresponding to the system to be analyzed.
[0076] In this step, threat factors in the threat factor set are transformed into manageable clue elements, and an initial clue weight value is assigned to each clue element. A multi-dimensional comprehensive evaluation method is used, which combines vendor risk, vulnerability severity, and developer behavior risk to assign initial clue weight values to clue elements, ensuring that the system can prioritize high-risk clues.
[0077] In one possible implementation of this application, step S102 may include: S1021. Obtain the correlation between each threat factor in the threat factor set.
[0078] For example, the correlation between each threat factor may include: vendor X provides hardware Y, and hardware Y has vulnerability Z, etc., that is, vendor to hardware to vulnerability; developer to software to vulnerability.
[0079] S1022. Based on the association, perform data mining on the set of threat factors to determine multiple clue elements composed of the threat factors, and use a preset graph database tool to determine the association and combination relationship between the clue elements.
[0080] In this application embodiment, the types of clue elements include hardware, software, suppliers, developers, and vulnerabilities, and each type corresponds to a different combination of threat factors based on the relationship.
[0081] For example, a clue element combining a vendor and hardware indicates that a hardware component provided by a vendor may be risky; a clue element combining a vulnerability and a developer indicates that a developer may have introduced vulnerable code.
[0082] S1023. For each of the clue elements, determine the clue weight value corresponding to the clue element based on the factor weight value corresponding to each of the threat factors corresponding to the clue element.
[0083] In this step, the specific implementation involves first obtaining the factor weight values corresponding to the supply chain layer threat factor and the developer threat factor in each clue element, namely, the supply chain layer factor weight value and the developer layer factor weight value; then, based on the hardware layer factor weight value and the software layer factor weight value, calculating the vulnerability weight value (vulnerability basic severity score); finally, combining the supply chain layer factor weight value, the developer layer factor weight value, and the vulnerability weight value, determining the clue weight value corresponding to each clue element.
[0084] In this embodiment of the application, the clue weight value corresponding to each clue element is determined by the following formula.
[0085] .
[0086] in, Indicates the first The clue weight value corresponding to each clue element; , , This represents the preset weighting coefficients (e.g., , , These can be set to 0.3, 0.1, or 0.6 respectively. This represents the weight values of factors at the supply chain level; This represents the developer-level factor weight value; This represents the vulnerability weight value calculated based on the hardware layer factor weight value and the software layer factor weight value.
[0087] Here, the weighting coefficient , , This reflects the relative importance of the three dimensions of vendors, developers, and vulnerabilities in the overall threat assessment. This coefficient allocation is based on the expert experience in the field of threat intelligence, with the vulnerability dimension having the highest weight because it is directly related to exploitable system weaknesses.
[0088] The clue weight value corresponding to each clue element enables the system to prioritize high-risk clues, thereby improving the efficiency and accuracy of threat analysis.
[0089] S1024. The clue elements, the association combination relationship, and the clue weight value corresponding to each clue element are determined as the clue element dataset corresponding to the system to be analyzed.
[0090] In this step, the clue elements, their associated combinations, and the clue weight value corresponding to each clue element are stored in the clue database, providing a data foundation for subsequent dynamic association analysis.
[0091] S103. Perform weighted correlation analysis on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset.
[0092] In this embodiment of the application, the weighted apriori with pruning and FP-growth algorithm is used to mine frequently occurring high-risk clue element groups from the clue element dataset. The W-APF algorithm is based on the improved weighted frequent pattern tree (FP-Tree) structure and combines multi-dimensional transaction weight calculation, which can effectively identify low-frequency high-risk threat combinations and make up for the shortcomings of traditional association rule mining.
[0093] The advantages of the W-APF algorithm are as follows: First, it can identify low-frequency high-risk threat combinations, making up for the shortcomings of the traditional Apriori algorithm which only relies on frequency (with a default weight of 1); second, it reflects the differences in importance of different threat factors through a weight mechanism, improving the accuracy of threat assessment; and finally, it supports dynamic weight updates, which can adapt to the dynamic characteristics of threat factors changing over time and in the environment.
[0094] Here, the target risk clue element group represents a combination of clue elements with a high risk of cyber threat.
[0095] In one possible implementation of this application, step S103 may include: S1031. Scan the clue element dataset to extract a first frequent itemset from the clue element dataset, and determine the clue weight value and association combination relationship corresponding to each clue element in the first frequent itemset.
[0096] In this step, the dataset of clue elements in the clue database is scanned, and the individual clue elements that appear most frequently and their clue weight values are extracted to obtain the first frequent itemset. Then, based on the first frequent itemset, the clue weight value and association combination relationship corresponding to each clue element are determined.
[0097] The first frequent itemset represents the frequent itemsets of the first scan of the clue element dataset to extract the individual clue elements that appear frequently and their clue weight values.
[0098] S1032. Arrange the clue elements in the first frequent itemset in descending order according to the numerical value of the clue weight to obtain the first frequent itemset sequence.
[0099] The first frequent item sequence represents the sequence of clue elements obtained by arranging all clue elements in the first frequent item set in descending order according to the numerical value of the clue weight.
[0100] S1033. Based on the preset root node, the first frequent item sequence, the clue weight value, and the association combination relationship between the clue elements in the first frequent item set, a frequent pattern tree corresponding to the system to be analyzed is constructed by recursion.
[0101] For example, first, a preset root node "root" is created and marked as "NULL"; then, for each clue element in the first frequent itemset, the clue element is sorted according to the order of the clue element sequence; then, the sorted list of clue elements is passed to the insert tree function; finally, the frequent pattern tree corresponding to the system to be analyzed is constructed recursively. Each node in the frequent pattern tree stores the name of the clue element, the corresponding clue weight value, and the sum of the clue weight values of all clue elements under that node.
[0102] S1034. Perform frequent pattern growth mining on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed.
[0103] In one possible implementation of this application, step S1034 may include: S10341. Determine whether there is only one target path in the frequent pattern tree.
[0104] In this step, it is determined whether there is only one target path in the frequent pattern tree based on the node structure of the frequent pattern tree.
[0105] S10342. If there is only one target path in the frequent pattern tree, then add the nodes corresponding to all clue elements in the frequent pattern tree to the target path, so as to determine the union of the frequent pattern tree and the first frequent itemset as the target frequent itemset.
[0106] For example, if there is only one path P in the frequent pattern tree β, then all nodes in β are added to path P to generate the union (βUα) of the frequent pattern tree and the first frequent itemset as the target frequent itemset.
[0107] S10343. If there are multiple target paths in the frequent pattern tree, then traverse the nodes corresponding to all clue elements in the frequent pattern tree to extract multiple conditional pattern bases related to the first frequent itemset, and calculate the weighted support corresponding to each conditional pattern base.
[0108] In this embodiment of the application, the weighted support corresponding to the conditional pattern base or clue element is calculated using the following formula.
[0109] .
[0110] in, Represents the conditional pattern base or clue element ( The corresponding weighted support; Represents the conditional pattern base or clue element ( The corresponding clue weight value; This represents the sum of the clue weights corresponding to each node in the frequent pattern tree.
[0111] S10344. Compare each weighted support with a preset support threshold, and filter out the conditional pattern bases corresponding to the weighted support that are greater than or equal to the support threshold, so as to obtain the target frequent itemset composed of the conditional pattern bases.
[0112] For example, extract the conditional pattern base from the frequent pattern tree β; traverse all nodes in the frequent pattern tree β and extract the conditional pattern base related to the first frequent itemset α; calculate the weighted support of the conditional pattern base; recursively call the FP span growth algorithm to generate the union (βUα) of the frequent pattern tree and the first frequent itemset as the target frequent itemset.
[0113] S10345. Calculate the target weighted support for each clue element in the target frequent itemset and the weighted confidence for each clue element group in the target frequent itemset.
[0114] In this embodiment of the application, the target weighted support corresponding to each clue element is calculated using the following formula.
[0115] .
[0116] in, Representing clue elements The corresponding target weighted support; Representing clue elements The corresponding clue weight value; This represents the sum of the clue weights corresponding to each node in the frequent pattern tree.
[0117] In this embodiment of the application, the weighted confidence level corresponding to each clue element group is calculated using the following formula.
[0118] .
[0119] in, Represents a group of clue elements The corresponding weighted confidence level; Indicates the presence of clue elements and clue elements The clue weight value; Indicates the presence of clue elements The weight value of the clue.
[0120] Here, weighted support reflects the importance of the combination of clue elements in the overall threat environment, rather than just the frequency of occurrence; weighted confidence reflects the strength of the association between the combinations of clue elements.
[0121] S10346. Based on the target weighted support and the weighted confidence, determine the target score value corresponding to each clue element group.
[0122] In this embodiment of the application, the target score value corresponding to each clue element group is calculated using the following formula.
[0123] .
[0124] in, This represents the target score value corresponding to each group of clue elements; This represents the average target-weighted support among all clue elements in each clue element group; This represents the weighted confidence level corresponding to each clue element group; This represents the average clue weight among all clue elements in each clue element group.
[0125] Here, the overall score takes into account support, confidence, and average weight, providing a comprehensive indicator for risk assessment of the combination of clue elements.
[0126] S10347. Based on the target score value, determine at least one target risk clue element group corresponding to the system to be analyzed in the target frequent item set.
[0127] In this step, each target score is compared with a preset score threshold, and the clue element group corresponding to the target score value that is greater than or equal to the score threshold in the target frequent item set is determined as the target risk clue element group corresponding to the system to be analyzed.
[0128] S104. Based on the target risk clue element group, generate a network threat reasoning report corresponding to the system to be analyzed, so as to analyze the network threat risk of the system to be analyzed.
[0129] In this embodiment of the application, a network threat inference report is generated based on the analysis results of the target risk clue element group, and the weights of the clue elements and knowledge graph nodes are dynamically updated. Specifically, a knowledge graph-based threat inference technology is used, combined with an attack-defense game model and machine learning algorithms, to generate a comprehensive and accurate threat report, and the knowledge graph nodes are updated in a combination of top-down and bottom-up approaches.
[0130] In this step, the threat reasoning report is generated based on knowledge graph query and analysis. Specifically, first, the query pattern for threat reasoning is defined, including attack paths, risk levels, and potential impacts; second, relevant entities and relationships are extracted from the knowledge graph; then, the query results are analyzed to identify potential attack paths and risk points; finally, a threat reasoning report is generated, including risk level assessment, potential attack path analysis, and risk mitigation recommendations.
[0131] In one possible implementation of this application, step S104 may include: S1041. Use a preset knowledge graph to extract threat risk association information from the target risk clue element group.
[0132] In this step, a pre-defined knowledge graph is used to extract relevant entities and relationships from each target risk clue element group to obtain threat risk association information, specifically including information such as associated vulnerabilities, attack techniques, suppliers, and developers.
[0133] S1042. Based on the threat risk association information, determine the attack path information and risk information corresponding to the system to be analyzed.
[0134] For example, based on the ATT&CK framework and the CAPEC attack pattern library, possible attack paths are analyzed; then, the risk level is assessed by combining vulnerability CVSS scores, vendor risk levels, and developer behavior risks, to obtain the attack path information and risk information corresponding to the system to be analyzed.
[0135] S1043. Based on the attack path information and the risk information, determine the mitigation suggestion information corresponding to the system to be analyzed.
[0136] For example, based on attack path information and risk information, combined with vulnerability remediation guidelines, vendor security requirements, and developer security specifications, mitigation recommendations are generated for the system to be analyzed.
[0137] S1044. Based on the attack path information, the risk information, and the mitigation suggestion information, generate a network threat inference report corresponding to the system to be analyzed.
[0138] For example, attack path information, risk information, and mitigation suggestion information are integrated into a network threat inference report corresponding to the system to be analyzed in a preset format (e.g., JSON format).
[0139] In another possible implementation of this application, step S104 further includes: S1045. In response to the user obtaining external intelligence data, based on the external intelligence data, the network threat reasoning report, and the target risk clue element group, the knowledge graph is updated and replaced to obtain an updated knowledge graph.
[0140] In this embodiment, the knowledge graph is updated using a combination of top-down and bottom-up approaches. Specifically, firstly, top-down updates adjust the entity types and relationships in the knowledge graph based on threat reasoning results and expert knowledge. For example, when a new attack technique is discovered, it may be necessary to add new entity types or relationship types. Secondly, bottom-up updates update the entity attributes and relationship weights in the knowledge graph based on new clue elements and association analysis results. For example, when a vulnerability is actually exploited, its CVSS score may need to be adjusted, thereby affecting its weight in the knowledge graph.
[0141] For example, as a knowledge graph update example: 1) Input: Threat reasoning report, new clue element, external intelligence; 2) Schema layer update: Adjust entity type and relation type as necessary; 3) Data layer update: Update entity attributes and relation weights via Cypher statements: A) Add new entity: "MERGE (v:Vulnerability {id:"CVE-2025-XXX'})", B) Update entity attributes: "SET v += {cvss:9.0, description:"..."}", C) Add / update relation: "MERGE (v)-[r:UsedIn {weight:0.7}]->(a:Attack)"; 4) Knowledge graph validation: Check whether the updated knowledge graph satisfies consistency constraints; 5) Output: Updated knowledge graph.
[0142] Optional, please refer to Figure 2 , Figure 2 This is a second flowchart illustrating a method for analyzing system network threat risks provided in an embodiment of this application. Figure 2 As shown in the figure, the system network threat risk analysis method provided in this application embodiment includes step S105 in addition to steps S101 to S104. Specifically, step S105 is used to realize the continuous evolution and closed-loop optimization of clue elements.
[0143] S105. In response to the system to be analyzed triggering at least one preset clue lifecycle management condition, perform the operation corresponding to the triggered clue lifecycle management condition on the system to be analyzed.
[0144] The conditions for managing the lifecycle of clues include conditions for clue generation, conditions for association activation, conditions for weight update, conditions for merging or downgrading clues, and conditions for knowledge accumulation.
[0145] In this step, the operations corresponding to the lead lifecycle management conditions include: lead generation, automatic triggering and initial weight allocation; association activation, dynamic activation and weight increase; weight update, periodic and event-driven weight adjustment; lead merging or downgrading, automatic and manual intervention and lead management; knowledge accumulation, accumulation after meeting the conditions to form a knowledge system.
[0146] In this way, the closed-loop management mechanism ensures that clue elements can be dynamically adjusted as the threat environment changes, while valuable clues are accumulated into knowledge, forming a continuously evolving knowledge system, thereby improving the efficiency and accuracy of threat analysis.
[0147] In this embodiment, the clue generation stage is the starting point of the lifecycle. Specifically: First, the triggering conditions for clue generation are defined, including vulnerability disclosure, vendor risk events, abnormal developer behavior, etc.; Second, when the triggering conditions are met, the system automatically generates new clue elements. For example, when a new CVE vulnerability is disclosed, the system automatically generates clue elements related to the vulnerability, including the vulnerability itself, potentially affected software, related vendors, etc.; Finally, initial weights are assigned to the new clue elements based on a comprehensive calculation of vendor weight, vulnerability weight, and developer weight.
[0148] The conditions for generating clues include: vulnerability disclosure, when a new CVE vulnerability is included in vulnerability databases such as NVD; vendor risk events, when a vendor experiences a security incident (such as a data breach or financial crisis); abnormal developer behavior, when abnormal developer behavior is detected (such as submitting outside of working hours or frequently modifying critical code); asset changes, when the system adds or removes assets such as hardware, software, or services; and periodic scanning, where the system periodically (e.g., weekly) scans all assets to generate new clue elements.
[0149] In this embodiment, the association activation phase is a key link in the clue lifecycle. Specifically, firstly, the conditions for association activation are defined, including clue elements being referenced by other targets, matching a known attack chain, and being related to the current security event. Secondly, when the association activation conditions are met, the system increases the weight of the clue element and marks it as "active". For example, when a vulnerability is discovered to be actually exploited, the weight of its associated clue elements will increase accordingly. Finally, the associated activated clue elements are included in the real-time monitoring scope to improve the targeting of threat analysis.
[0150] The associated activation conditions include: being referenced by other targets (the clue element is referenced by multiple target systems); matching attack chains (the clue element matches a known attack chain, such as the ATT&CK framework); being related to security events (the clue element is related to the currently detected security event); and a weight threshold (the weight of the clue element exceeds a certain threshold, such as 0.5).
[0151] In this embodiment, the weight update phase is the core of the clue lifecycle. Specifically, firstly, the triggering conditions for weight updates are defined, including vulnerability patching, vendor risk mitigation, and normalization of developer behavior. Secondly, when the triggering conditions are met, the system recalculates the weight of the clue element according to the updated weight calculation formula. For example, when a vulnerability is patched, its weight will decrease accordingly, and when vendor risk is mitigated, the weight of its associated clue element will also decrease. Finally, the updated weights are stored in the clue database and reflected in the knowledge graph.
[0152] The triggering conditions for weight updates are as follows: vulnerability patching, when a vulnerability is patched or mitigation measures are implemented; vendor risk mitigation, when the vendor risk score decreases; normalization of developer behavior, when developer behavior returns to normal; time decay, as the weight of the clue element gradually decays over time; and expert intervention, when security experts adjust the weight of the clue element.
[0153] In this embodiment of the application, the dynamic adjustment of vulnerability weights is based on an attack-defense game model and a machine learning algorithm. Specifically, firstly, an attack-defense game model is defined, including the strategy sets of the attacker and the defender, utility functions, etc.; secondly, the attack cost and defense cost of the vulnerability are calculated through the game model, and then the vulnerability weights are adjusted.
[0154] In this embodiment of the application, the vulnerability weight is updated through the following steps.
[0155] “ " 。
[0156] Furthermore, the attack cost is calculated using the CVSS availability formula, as shown below.
[0157] .
[0158] in, , , and These represent CVSS metrics such as attack vector, attack complexity, required permissions, and user interaction, respectively. This indicates the cost of the attack.
[0159] Furthermore, defense costs are calculated by analyzing an organization's defense measures (such as firewalls, intrusion detection systems, security training, etc.).
[0160] Furthermore, by incorporating the time decay factor, the vulnerability weights are further adjusted; that is, the product of the time decay factor and the vulnerability weight is determined as the updated vulnerability weight.
[0161] The time decay factor is calculated using the gamma distribution function, and its specific representation is shown below.
[0162] .
[0163] in, Indicates the first The time decay factor is updated next time; This is the input parameter, usually set to 5. Setting the input parameter to 5 means that the weight is higher in the first 5 days after the vulnerability is disclosed, and then gradually decreases thereafter.
[0164] For example, when a vulnerability is actually exploited, the vulnerability weight is increased (e.g., CVSS score × 1.2); when a vulnerability is patched, the vulnerability weight is decreased (e.g., CVSS score × 0.4); when a vendor experiences a security incident, the weight of its associated clues is increased (e.g., vendor weight × 1.5); when a developer behaves abnormally, the weight of its associated clues is increased (e.g., developer weight × 1.3); and a time decay factor is applied periodically (e.g., monthly) to adjust the weight of all clues.
[0165] In this embodiment, the lead merging or downgrading phase is a management stage in the lead lifecycle. Specifically, firstly, the conditions for lead merging are defined, including the same lead element type, the same CVE ID, the same supplier, etc.; secondly, when the merging conditions are met, the system merges the lead elements using Neo4j's APOC function, i.e., "CALLapoc.refactor.mergeNodes([v1, v2], {properties:"keep", merge relationships:"true"})"; then, the conditions for lead downgrading are defined, including the lead element being unrelated for a long time (e.g., not activated for more than 6 months), having a weight below a threshold (e.g., 0.2), and having a vulnerability patched and no new exploits; finally, when the downgrading conditions are met, the system marks the lead element as "low risk" or "resolved" and removes it from the high-risk monitoring list.
[0166] The conditions for merging or downgrading leads include: lead merging, which means that the CVE IDs are the same, the vendors are the same and the risks are the same, the developers are the same and the behavior patterns are the same; lead downgrading, which means that the weight is less than 0.2, there is no long-term correlation (more than 6 months), the vulnerability has been fixed and there are no new exploits, the vendor risk has been mitigated, and the developer's behavior has returned to normal.
[0167] In this embodiment, the knowledge accumulation stage is the end of the clue lifecycle. Specifically, firstly, the conditions for knowledge accumulation are defined, including clue element weight exceeding a threshold (e.g., 0.8), multiple activations, and high correlation with known attack chains. Secondly, when the accumulation conditions are met, the system converts the clue element into a knowledge node in the knowledge graph and establishes associations with entities such as attack techniques and APT organizations. For example, when a vulnerability is exploited multiple times and forms an attack chain, the system will accumulate it as a "known attack path" node in the knowledge graph. Finally, the accumulated knowledge nodes are incorporated into the threat intelligence database to provide a reference for future threat analysis.
[0168] The knowledge accumulation conditions include: weight threshold (the weight of the clue element exceeds 0.8); activation frequency (the clue element is activated multiple times, such as more than 5 times); association strength (the clue element has a high association strength with known attack chains); and expert verification (the clue element has been verified and confirmed by security experts).
[0169] The system network threat risk analysis method provided in this application decomposes the network threat risks existing in the system to be analyzed into multi-dimensional threat factors, generates a clue element dataset including clue elements and their weights, and performs weighted correlation analysis on the clue element dataset to mine target risk clue element groups. Then, based on the target risk clue element groups, a network threat reasoning report corresponding to the system to be analyzed is generated, realizing the proactive identification and evolution tracking of complex network threats. While deeply analyzing the components of the system and their interrelationships, it effectively responds to new threats such as advanced persistent threats and supply chain attacks, improving the accuracy and real-time performance of system network security situation awareness.
[0170] Please see Figure 3 , Figure 4 , Figure 3 This is one of the structural schematic diagrams of a system network threat risk analysis device provided in an embodiment of this application. Figure 4 This is a second schematic diagram of a system network threat risk analysis device provided in an embodiment of this application. Figure 3 As shown, the analysis device 300 includes: Threat factor splitting module 310 is used to obtain system information corresponding to the system to be analyzed, and based on the system information, split and weight the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed. The clue element generation module 320 is used to determine the clue element dataset corresponding to the system to be analyzed based on the threat factor set and the correlation between each threat factor in the threat factor set; The weighted correlation analysis module 330 is used to perform weighted correlation analysis on the clue element dataset and determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset. The reasoning report generation module 340 is used to generate a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group, so as to analyze the network threat risk of the system to be analyzed.
[0171] Furthermore, when the threat factor splitting module 310 is used to acquire system information corresponding to the system to be analyzed, and based on the system information, to split and weight the network threats of the system to be analyzed, and determine the set of threat factors corresponding to the system to be analyzed, the threat factor splitting module 310 is used to: Obtain system information corresponding to the system to be analyzed; wherein, the system information includes at least hardware component information, software information, supplier information, and developer behavior information; Based on the system information, the network threats of the system to be analyzed are split into threat factors, and the hardware layer threat factors, software layer threat factors, supply chain layer threat factors and developer layer threat factors corresponding to the system to be analyzed are determined respectively. Based on the hardware score value corresponding to the hardware component information and multiple preset hardware factor coefficients, the hardware layer factor weight value corresponding to the hardware layer threat factor is determined. Based on the software score corresponding to the software information and multiple preset software factor coefficients, determine the software layer factor weight value corresponding to the software layer threat factor; Based on the supplier rating value corresponding to the supplier information and multiple preset supply chain factor coefficients, the supply chain layer factor weight value corresponding to the supply chain layer threat factor is determined. Based on the developer rating value corresponding to the developer behavior information and multiple preset developer factor coefficients, the developer layer factor weight value corresponding to the developer layer threat factor is determined. The hardware layer threat factor, the hardware layer factor weight value, the software layer threat factor, the software layer factor weight value, the supply chain layer threat factor, the supply chain layer factor weight value, the developer layer threat factor, and the developer layer factor weight value are determined as the set of threat factors corresponding to the system to be analyzed.
[0172] Furthermore, when the clue element generation module 320 determines the clue element dataset corresponding to the system to be analyzed based on the threat factor set and the correlation between each threat factor in the threat factor set, the clue element generation module 320 is used to: Obtain the correlation between each threat factor in the threat factor set; Based on the aforementioned relationships, data mining is performed on the set of threat factors to identify multiple clue elements composed of the threat factors, and a preset graph database tool is used to determine the association and combination relationships between the clue elements. For each of the aforementioned clue elements, a clue weight value corresponding to that clue element is determined based on the factor weight value corresponding to each of the aforementioned threat factors. The clue elements, the associated combination relationships, and the clue weight value corresponding to each clue element are determined as the clue element dataset corresponding to the system to be analyzed.
[0173] Furthermore, when the weighted correlation analysis module 330 performs weighted correlation analysis on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset, the weighted correlation analysis module 330 is used to: The dataset of clue elements is scanned to extract a first frequent itemset from the dataset, and the clue weight value and association combination relationship corresponding to each clue element in the first frequent itemset are determined. The clue elements in the first frequent itemset are sorted in descending order according to the numerical value of the clue weight to obtain the first frequent itemset sequence; Based on the pre-defined root node, the first frequent item sequence, the cue weight value, and the association combination relationship between the cue elements in the first frequent item set, a frequent pattern tree corresponding to the system to be analyzed is constructed by recursion. Frequent pattern growth mining is performed on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed.
[0174] Furthermore, when the weighted association analysis module 330 performs frequent pattern growth mining on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed, the weighted association analysis module 330 is used to: Determine whether there is only one target path in the frequent pattern tree; If there is only one target path in the frequent pattern tree, then the nodes corresponding to all clue elements in the frequent pattern tree are added to the target path, so that the union of the frequent pattern tree and the first frequent itemset is determined as the target frequent itemset. If there are multiple target paths in the frequent pattern tree, then traverse the nodes corresponding to all clue elements in the frequent pattern tree to extract multiple conditional pattern bases related to the first frequent itemset, and calculate the weighted support corresponding to each conditional pattern base. Each weighted support is compared with a preset support threshold, and the conditional pattern bases corresponding to the weighted support that are greater than or equal to the support threshold are selected to obtain the target frequent itemset composed of the conditional pattern bases. Calculate the target weighted support for each clue element in the target frequent itemset, and the weighted confidence for each group of clue elements in the target frequent itemset. Based on the target weighted support and the weighted confidence, determine the target score value corresponding to each clue element group; Based on the target score, at least one group of target risk clue elements corresponding to the system to be analyzed is determined from the target frequent itemset.
[0175] Furthermore, when the reasoning report generation module 340 generates a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group, the reasoning report generation module 340 is used to: Threat and risk correlation information is extracted from the target risk clue element group using a pre-defined knowledge graph. Based on the threat risk association information, the attack path information and risk information corresponding to the system to be analyzed are determined respectively; Based on the attack path information and the risk information, mitigation suggestions are determined for the system to be analyzed. Based on the attack path information, the risk information, and the mitigation suggestion information, a network threat inference report corresponding to the system to be analyzed is generated.
[0176] Furthermore, when generating a network threat inference report corresponding to the system to be analyzed based on the target risk clue element group, the inference report generation module 340 is also used for: In response to a user obtaining external intelligence data, the knowledge graph is updated and replaced based on the external intelligence data, the network threat reasoning report, and the target risk clue element group to obtain an updated knowledge graph.
[0177] Furthermore, such as Figure 4 As shown, the analysis device 300 further includes a lifecycle management module 350, which is used for: In response to the system under analysis triggering at least one preset clue lifecycle management condition, the system under analysis performs the operation corresponding to the triggered clue lifecycle management condition; wherein, the clue lifecycle management condition includes clue generation condition, association activation condition, weight update condition, clue merging or downgrading condition, and knowledge accumulation condition.
[0178] The system network threat risk analysis device provided in this application decomposes the network threat risks existing in the system to be analyzed into multi-dimensional threat factors, generates a clue element dataset including clue elements and their weights, and performs weighted correlation analysis on the clue element dataset to mine target risk clue element groups. Then, based on the target risk clue element groups, a network threat reasoning report corresponding to the system to be analyzed is generated, realizing the proactive identification and evolution tracking of complex network threats. While deeply analyzing the components of the system and their interrelationships, it effectively responds to new threats such as advanced persistent threats and supply chain attacks, improving the accuracy and real-time performance of system network security situation awareness.
[0179] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.
[0180] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 is running, the processor 510 and the memory 520 communicate via the bus 530. When the machine-readable instructions are executed by the processor 510, they can perform the operations described above. Figure 1 as well as Figure 2 The steps of the system network threat risk analysis method in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0181] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 as well as Figure 2 The steps of the system network threat risk analysis method in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0183] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0184] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0185] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0186] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0187] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for analyzing system network threat risks, characterized in that, The analytical method includes: Obtain system information corresponding to the system to be analyzed, and based on the system information, decompose and weight the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed; Based on the set of threat factors and the correlation between each threat factor in the set of threat factors, the dataset of clue elements corresponding to the system to be analyzed is determined. A weighted association analysis is performed on the clue element dataset to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset; Based on the target risk clue element group, a network threat reasoning report corresponding to the system to be analyzed is generated to analyze the network threat risk of the system to be analyzed.
2. The method according to claim 1, characterized in that, The process of acquiring system information corresponding to the system to be analyzed, and based on the system information, decomposing and weighting the network threats of the system to be analyzed to determine the set of threat factors corresponding to the system to be analyzed, includes: Obtain system information corresponding to the system to be analyzed; wherein, the system information includes at least hardware component information, software information, supplier information, and developer behavior information; Based on the system information, the network threats to the system to be analyzed are split into threat factors, and the hardware layer threat factors, software layer threat factors, supply chain layer threat factors and developer layer threat factors corresponding to the system to be analyzed are determined respectively. Based on the hardware score value corresponding to the hardware component information and multiple preset hardware factor coefficients, the hardware layer factor weight value corresponding to the hardware layer threat factor is determined. Based on the software score corresponding to the software information and multiple preset software factor coefficients, determine the software layer factor weight value corresponding to the software layer threat factor; Based on the supplier rating value corresponding to the supplier information and multiple preset supply chain factor coefficients, the supply chain layer factor weight value corresponding to the supply chain layer threat factor is determined. Based on the developer rating value corresponding to the developer behavior information and multiple preset developer factor coefficients, the developer layer factor weight value corresponding to the developer layer threat factor is determined. The hardware layer threat factor, the hardware layer factor weight value, the software layer threat factor, the software layer factor weight value, the supply chain layer threat factor, the supply chain layer factor weight value, the developer layer threat factor, and the developer layer factor weight value are determined as the set of threat factors corresponding to the system to be analyzed.
3. The method according to claim 1, characterized in that, The step of determining the cue element dataset corresponding to the system to be analyzed based on the set of threat factors and the correlation between each threat factor in the set of threat factors includes: Obtain the correlation between each threat factor in the threat factor set; Based on the aforementioned relationships, data mining is performed on the set of threat factors to identify multiple clue elements composed of the threat factors, and a preset graph database tool is used to determine the association and combination relationships between the clue elements. For each of the aforementioned clue elements, a clue weight value corresponding to that clue element is determined based on the factor weight value corresponding to each of the aforementioned threat factors. The clue elements, the associated combination relationships, and the clue weight value corresponding to each clue element are determined as the clue element dataset corresponding to the system to be analyzed.
4. The method according to claim 1, characterized in that, The weighted association analysis of the clue element dataset, to determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset, includes: The dataset of clue elements is scanned to extract a first frequent itemset from the dataset, and the clue weight value and association combination relationship corresponding to each clue element in the first frequent itemset are determined. The clue elements in the first frequent itemset are sorted in descending order according to the numerical value of the clue weight to obtain the first frequent itemset sequence; Based on the pre-defined root node, the first frequent item sequence, the cue weight value, and the association combination relationship between the cue elements in the first frequent item set, a frequent pattern tree corresponding to the system to be analyzed is constructed by recursion. Frequent pattern growth mining is performed on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed.
5. The method according to claim 4, characterized in that, The step of performing frequent pattern growth mining on the frequent pattern tree to determine at least one target risk clue element group corresponding to the system to be analyzed includes: Determine whether there is only one target path in the frequent pattern tree; If there is only one target path in the frequent pattern tree, then the nodes corresponding to all clue elements in the frequent pattern tree are added to the target path, so that the union of the frequent pattern tree and the first frequent itemset is determined as the target frequent itemset. If there are multiple target paths in the frequent pattern tree, then traverse the nodes corresponding to all clue elements in the frequent pattern tree to extract multiple conditional pattern bases related to the first frequent itemset, and calculate the weighted support corresponding to each conditional pattern base. Each weighted support is compared with a preset support threshold, and the conditional pattern bases corresponding to the weighted support that are greater than or equal to the support threshold are selected to obtain the target frequent itemset composed of the conditional pattern bases. Calculate the target weighted support for each clue element in the target frequent itemset, and the weighted confidence for each group of clue elements in the target frequent itemset. Based on the target weighted support and the weighted confidence, determine the target score value corresponding to each clue element group; Based on the target score, at least one group of target risk clue elements corresponding to the system to be analyzed is determined from the target frequent itemset.
6. The method according to claim 1, characterized in that, The process of generating a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group includes: Threat and risk correlation information is extracted from the target risk clue element group using a pre-defined knowledge graph. Based on the threat risk association information, the attack path information and risk information corresponding to the system to be analyzed are determined respectively; Based on the attack path information and the risk information, mitigation suggestions are determined for the system to be analyzed. Based on the attack path information, the risk information, and the mitigation suggestion information, a network threat inference report corresponding to the system to be analyzed is generated.
7. The method according to claim 6, characterized in that, The step of generating a network threat reasoning report corresponding to the system to be analyzed based on the target risk clue element group also includes: In response to a user obtaining external intelligence data, the knowledge graph is updated and replaced based on the external intelligence data, the network threat reasoning report, and the target risk clue element group to obtain an updated knowledge graph.
8. The method according to claim 1, characterized in that, After generating the network threat inference report corresponding to the system to be analyzed, the analysis method further includes: In response to the system under analysis triggering at least one preset clue lifecycle management condition, the system under analysis performs the operation corresponding to the triggered clue lifecycle management condition; wherein, the clue lifecycle management condition includes clue generation condition, association activation condition, weight update condition, clue merging or downgrading condition, and knowledge accumulation condition.
9. A system network threat risk analysis device, characterized in that, The analytical device includes: The threat factor splitting module is used to obtain system information corresponding to the system to be analyzed, and based on the system information, to split and weight the network threats of the system to be analyzed, and determine the set of threat factors corresponding to the system to be analyzed. The clue element generation module is used to determine the clue element dataset corresponding to the system to be analyzed based on the threat factor set and the correlation between each threat factor in the threat factor set; The weighted correlation analysis module is used to perform weighted correlation analysis on the clue element dataset and determine at least one target risk clue element group corresponding to the system to be analyzed in the clue element dataset. The inference report generation module is used to generate a network threat inference report corresponding to the system to be analyzed based on the target risk clue element group, so as to analyze the network threat risk of the system to be analyzed.
10. An electronic device, characterized in that, include: The system includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the system network threat risk analysis method as described in any one of claims 1 to 7.