A knowledge graph-based method and system for linking ATT&CK and CVE
Patent Information
- Application Number
- CN202311354104.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-18
AI Technical Summary
在当安全专业人员对于相关系统漏洞相关知识体系掌握不全面时,安全专业人员不能准确地分析出攻击者采用的攻击手段所具体针对的系统漏洞,进而导致系统的遭受威胁的概率和潜在风险增加
1、首先基于ATT&CK模型框架和多个预设数据集构建知识图谱,知识图谱中包括多个CVE实体、多个CWE实体以及多个ATT&CK实体,从而能够全面地反映出各个实体之间的关系;再通过将各个第二CVE实体和第二CVE实体对应的实体属性输入至第一文本分类模型,输出得到与第二CVE实体对应的CWE编号,能够准确地为没有明确对应CWE编号的CVE实体找到对应的CWE编号;再基于知识图谱和图论相关算法,对CVE实体集中全部CVE实体进行分析,构建全部CVE实体与全部ATT&CK实体的关联映射关系,能够清晰地展示出CVE系统漏洞和ATT&CK攻击技术之间的关联关系,从而可以提升分析出攻击者采用的攻击手段所具体针对的系统漏洞的准确程度,进而降低系统的遭受威胁的概率和潜在风险。
Smart Images

Figure CN117407884B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, specifically to a method and system for associating ATT&CK and CVE based on knowledge graphs. Background Technology
[0002] With the rapid development and widespread application of information technology, cybersecurity threats are increasing. Existing cybersecurity analysis methods often employ the comprehensive theoretical attack knowledge system provided by the ATT&CK model framework to help security professionals understand attackers' thought processes and attack patterns.
[0003] However, attackers typically target system vulnerabilities when attacking a system. When security professionals lack a comprehensive understanding of the relevant system vulnerabilities, they cannot accurately analyze the specific vulnerabilities targeted by the attacker's methods, thus increasing the probability of the system being threatened and the potential risks.
[0004] Therefore, there is an urgent need for a knowledge graph-based method and system for associating ATT&CK and CVE to solve the problems existing in the above-mentioned background technology. Summary of the Invention
[0005] This application provides a knowledge graph-based method and system for associating ATT&CK and CVE, which can improve the accuracy of analyzing the specific system vulnerabilities targeted by the attacker's attack methods, thereby reducing the probability of the system being threatened and the potential risks.
[0006] Firstly, this application provides a knowledge graph-based method for associating ATT&CK and CVE, the method comprising: constructing a knowledge graph based on the ATT&CK model framework and multiple preset datasets, wherein the knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities; the multiple CVE entities include first CVE entities and second CVE entities; the entity attributes of the CWE entities include CWE numbers, and the first CVE entities have explicitly corresponding CWE numbers, while the second CVE entities do not have explicitly corresponding CWE numbers; inputting each second CVE entity and its corresponding entity attributes into a first text classification model, and outputting the CWE numbers corresponding to the second CVE entities; analyzing all CVE entities in the CVE entity set based on the knowledge graph and graph theory-related algorithms, and constructing an association mapping relationship between all CVE entities and all ATT&CK entities; wherein the CVE entity set includes multiple first CVE entities and multiple second CVE entities; and outputting the association mapping relationship.
[0007] By adopting the above technical solution, a knowledge graph is first constructed based on the ATT&CK model framework and multiple preset datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities, thus comprehensively reflecting the relationships between various entities. Next, by inputting the entity attributes corresponding to each second CVE entity into the first text classification model, the CWE number corresponding to the second CVE entity is output, accurately finding the corresponding CWE number for CVE entities without a clearly defined CWE number. Then, based on the knowledge graph and graph theory-related algorithms, all CVE entities in the CVE entity set are analyzed, constructing an association mapping relationship between all CVE entities and all ATT&CK entities. This clearly demonstrates the correlation between CVE system vulnerabilities and ATT&CK attack techniques, thereby improving the accuracy of analyzing the specific system vulnerabilities targeted by the attacker's attack methods, and thus reducing the probability of system threats and potential risks.
[0008] Optionally, the knowledge graph further includes multiple CAPEC entities; the association mapping relationship includes multiple first sub-mapping relationships and multiple second sub-mapping relationships; the step of analyzing all CVE entities in the CVE entity set based on the knowledge graph and graph theory-related algorithms, and constructing the association mapping relationship between all CVE entities and all ATT&CK entities, specifically includes: determining whether the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity; when the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity, determining whether the current CAPEC entity and the ATT&CK entity exist. In the mapping relationship; when the current CAPEC entity has a mapping relationship with the ATT&CK entity, based on the knowledge graph, a first sub-mapping relationship is obtained between the CVE entity corresponding to the current CAPEC entity and the ATT&CK entity; when the current CAPEC entity does not have a mapping relationship with the ATT&CK entity, based on the CAPEC entity set corresponding to the current CWE entity, the K first CAPEC entities with the highest similarity to the current CAPEC entity are selected from the CAPEC entity set; based on the knowledge graph, a second sub-mapping relationship is obtained between the CVE entities corresponding to the K first CAPEC entities and the ATT&CK entity.
[0009] Optionally, the entity attributes of the CWE entity further include first text information; after determining whether the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity, the method further includes: when the current CWE entity corresponding to each CVE entity is not mapped to any CAPEC entity, constructing an association set of first CWE entities associated with the vulnerability type of the current CWE entity; when the association set is a non-empty set, obtaining at least one CAPEC entity associated with the current CWE entity based on the mapping relationship between the first CWE entity and the CAPEC entity; when the association set is an empty set, inputting the first text information into a second text classification model to obtain a second CWE entity; the similarity between the first text information corresponding to the second CWE entity and the first text information corresponding to the current CWE entity is greater than a first preset similarity threshold; obtaining at least one CAPEC entity associated with the current CWE entity based on the mapping relationship between the second CWE entity and the CAPEC entity.
[0010] By adopting the above technical solution, by determining whether the CWE entity corresponding to the CVE entity is mapped to at least one CAPEC entity, and whether there is a mapping relationship between the CAPEC entity and the ATT&CK entity, the association relationship between the CVE entity, CWE entity, CAPEC entity and ATT&CK entity can be more deeply explored. When there is no mapping relationship between the CAPEC entity and the ATT&CK entity, based on the CAPEC entity set corresponding to the current CWE entity, the K first CAPEC entities with the highest similarity to the current CAPEC entity are selected from the CAPEC entity set, and the mapping relationship between the CVE entity and the ATT&CK entity is further constructed, thereby improving the accuracy of the mapping relationship.
[0011] Optionally, the step of selecting the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set corresponding to the current CWE entity specifically includes: calculating the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity; selecting L second CAPEC entities from the CAPEC entity set; the Jaccard coefficient between the second CAPEC entity and the current CAPEC entity is greater than a preset Jaccard coefficient threshold; and filtering and converging the L second CAPEC entities based on the n-order association relationship corresponding to the second CAPEC entities to obtain K first CAPEC entities, where n is a positive integer greater than 1.
[0012] By adopting the above technical solution, and by calculating the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity, it is possible to accurately filter out the second CAPEC entity that is more similar to the current CAPEC entity. Based on the n-order association relationship corresponding to the second CAPEC entity, the second CAPEC entity is filtered and converged. This allows for in-depth mining of the association relationships between CVE entities, CWE entities, CAPEC entities, and ATT&CK entities, thereby improving the flexibility and accuracy of the mapping and further enhancing the accuracy of the relationship analysis between CVE entities and ATT&CK entities.
[0013] Optionally, calculating the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity specifically includes: calculating the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity using the following formula: Where A is the set of elements in the current CAPEC entity, B is the set of elements in any CAPEC entity in the CAPEC entity set, || is the number of elements in the set, ∩ is the intersection, and ∪ is the union.
[0014] Optionally, the K first CAPEC entities include X directly associated CAPEC entities and Y indirectly associated CAPEC entities, where X + Y = K; the second sub-mapping relationship includes direct sub-mapping relationships and indirect sub-mapping relationships; obtaining the second sub-mapping relationship between the CVE entities corresponding to the K first CAPEC entities and the ATT&CK entities based on the knowledge graph specifically includes: obtaining the direct sub-mapping relationship based on the mapping relationship between the X directly associated CAPEC entities and the ATT&CK entities; inputting the second text information corresponding to the Y indirectly associated CAPEC entities and the third text information corresponding to the ATT&CK entities into a third text classification model, and calculating the cosine similarity between the second text information and the third text information; selecting multiple associated ATT&CK entities from all the ATT&CK entities, wherein the cosine similarity between the third text information corresponding to each associated ATT&CK entity and the second text information corresponding to the indirectly associated CAPEC entity is greater than a second preset similarity threshold; obtaining the indirect sub-mapping relationship based on the mapping relationship between the Y indirectly associated CAPEC entities and the associated ATT&CK entities.
[0015] By adopting the above technical solution, the second text information corresponding to the indirectly associated CAPEC entity and the third text information corresponding to the ATT&CK entity are input into the third text classification model, and the cosine similarity between the second text information and the third text information is calculated, so as to more accurately filter out the ATT&CK entity that is highly associated with the current CAPEC entity.
[0016] Optionally, calculating the cosine similarity between the second text information and the third text information specifically includes: calculating the cosine similarity between the second text information and the third text information using the following formula: in, This is the embedding vector of the second text information. The embedding vector of the third text information. Let C be the magnitude of vector C. Let be the magnitude of vector D.
[0017] A second aspect of this application provides a knowledge graph-based association system for ATT&CK and CVE, the system comprising: a knowledge graph construction module, a processing module, and an output module; the knowledge graph construction module is used to construct a knowledge graph based on the ATT&CK model framework and multiple preset datasets, the knowledge graph including multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities; the multiple CVE entities include a first CVE entity and a second CVE entity; the entity attributes of the CWE entities include a CWE number, and the first CVE entity has a clearly corresponding CWE number, The second CVE entity does not have a clearly corresponding CWE number; the processing module is used to input each second CVE entity into the first text classification model and output the CWE number corresponding to the second CVE entity; the processing module is also used to analyze all CVE entities in the CVE entity set based on the knowledge graph and graph theory related algorithms, and construct the association mapping relationship between all the CVE entities and all the ATT&CK entities; wherein, the CVE entity set includes multiple first CVE entities and multiple second CVE entities; the output module is used to output the association mapping relationship.
[0018] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of the first aspects of this application.
[0019] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform a computer program as described in any of the first aspects of this application.
[0020] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. First, a knowledge graph is constructed based on the ATT&CK model framework and multiple pre-set datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities, thus comprehensively reflecting the relationships between various entities. Then, by inputting the entity attributes corresponding to each second CVE entity into the first text classification model, the CWE number corresponding to the second CVE entity is output, accurately finding the corresponding CWE number for CVE entities without a clearly defined CWE number. Next, based on the knowledge graph and graph theory-related algorithms, all CVE entities in the CVE entity set are analyzed, constructing an association mapping relationship between all CVE entities and all ATT&CK entities. This clearly demonstrates the correlation between CVE system vulnerabilities and ATT&CK attack techniques, thereby improving the accuracy of analyzing the specific system vulnerabilities targeted by attackers' attack methods, and ultimately reducing the probability of system threats and potential risks.
[0021] 2. By determining whether the CWE entity corresponding to the CVE entity maps to at least one CAPEC entity, and whether there is a mapping relationship between the CAPEC entity and the ATT&CK entity, the association relationships between CVE entities, CWE entities, CAPEC entities, and ATT&CK entities can be explored more deeply. When there is no mapping relationship between the CAPEC entity and the ATT&CK entity, based on the CAPEC entity set corresponding to the current CWE entity, the K first CAPEC entities with the highest similarity to the current CAPEC entity are selected from the CAPEC entity set, and the mapping relationship between the CVE entity and the ATT&CK entity is further constructed, thereby improving the accuracy of the mapping relationship.
[0022] 3. By calculating the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity, it is possible to accurately filter out the second CAPEC entity that is more similar to the current CAPEC entity. Based on the n-order association relationship corresponding to the second CAPEC entity, the second CAPEC entity is filtered and converged. This allows for in-depth mining of the association relationships between CVE entities, CWE entities, CAPEC entities, and ATT&CK entities, thereby improving the flexibility and accuracy of the mapping and further enhancing the accuracy of the relationship analysis between CVE entities and ATT&CK entities.
[0023] 4. By inputting the second text information corresponding to the indirectly associated CAPEC entity and the third text information corresponding to the ATT&CK entity into the third text classification model, the cosine similarity between the second and third text information is calculated, thereby enabling more accurate filtering of ATT&CK entities that are highly associated with the current CAPEC entity. Attached Figure Description
[0024] Figure 1 This is one of the flowcharts illustrating a knowledge graph-based association method for ATT&CK and CVE provided in this application embodiment; Figure 2 This is a schematic diagram illustrating entities, relationships, and attributes provided in an embodiment of this application; Figure 3 This is a second flowchart illustrating a knowledge graph-based association method for ATT&CK and CVE disclosed in an embodiment of this application. Figure 4 This is the third schematic diagram of the structure of a knowledge graph-based association system for ATT&CK and CVE disclosed in this application. Figure 5 This is the fourth schematic diagram of the structure of a knowledge graph-based association system for ATT&CK and CVE disclosed in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a knowledge graph-based association system for ATT&CK and CVE disclosed in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.
[0025] Explanation of reference numerals in the attached diagram: 1. Knowledge graph construction module; 2. Processing module; 3. Output module; 700. Electronic device; 701. Processor; 702. Communication bus; 703. User interface; 704. Network interface; 705. Memory. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0027] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0028] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0029] ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) is a knowledge base and model for attack behavior. It describes the techniques used at each stage of an attack from the attacker's perspective. The goal of ATT&CK is to create a comprehensive list of known adversarial tactics and techniques used in cyberattacks. The ATT&CK model is primarily used in areas such as assessing offensive and defensive capability coverage, APT intelligence analysis, threat hunting, and attack simulation. It details how each offensive and defensive technique is utilized, greatly helping security professionals quickly understand the relevant technical content. Within the ATT&CK technology matrix, each technique includes specific scenario examples illustrating how attackers utilize that technique through a particular piece of malware or action plan.
[0030] CVE (Common Vulnerabilities and Exposures) is a vulnerability database. Each vulnerability has a unique number; for example, the infamous log4j vulnerability is numbered CVE-2021-44228. The numbering system consists of the letters "CVE," the year of disclosure, and a 4-5 digit number. Vulnerabilities cover various aspects of computer systems, including protocol implementation vulnerabilities, software implementation vulnerabilities, and hardware vulnerabilities. CVE records are very brief, only providing a simple description of the vulnerability. More detailed information about vulnerabilities is collected in other databases, including the U.S. National Vulnerability Database (NVD), the China National Vulnerability Database (CNVD), the CERT / CC Vulnerability Annotation Database, and various lists maintained by vendors and other organizations. Using the CVE number, users can easily identify the same security flaw across these different systems.
[0031] Since attackers typically target system vulnerabilities when attacking a system, this application aims to combine the attack techniques within the ATT&CK model framework with the system vulnerability-related techniques in the CVE database. This will allow for the exploration of the correlation between specific ATT&CK attack techniques and CVE vulnerabilities, leading to a better understanding of the connection between different attack techniques and known vulnerabilities. This, in turn, will enable a better assessment of the probability of system threats and potential risks, thereby effectively improving system security.
[0032] Therefore, in order to combine the attack techniques related to the ATT&CK model framework with the system vulnerability-related technical content in the CVE database, and to explore the correlation between specific ATT&CK attack techniques and CVE vulnerabilities, this application provides a knowledge graph-based method for associating ATT&CK and CVE, referring to... Figure 1 , Figure 1 This is one of the flowcharts illustrating a knowledge graph-based association method for ATT&CK and CVE disclosed in this application. The knowledge graph-based association method for ATT&CK and CVE includes steps S1 to S4, as follows: Step S1: Construct a knowledge graph based on the ATT&CK model framework and multiple preset datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities. The multiple CVE entities include a first CVE entity and a second CVE entity. The entity attributes of the CWE entities include a CWE number, and the first CVE entity has a clearly corresponding CWE number, while the second CVE entity does not have a clearly corresponding CWE number.
[0033] In the steps described above, the server constructs a knowledge graph based on the ATT&CK model framework and multiple pre-set datasets.
[0034] Specifically, in this technical solution, multiple preset datasets include, but are not limited to, the ATT&CK dataset, CVE dataset, CAPEC dataset, etc. contained in the ATT&CK model framework.
[0035] The server first collects datasets such as ATT&CK, CVE, and CAPEC. Next, it translates the attributes and descriptions in the collected datasets into the required language, such as Chinese. Then, it cleans and removes duplicates from the collected datasets. Finally, it identifies relevant entities in the preprocessed datasets, such as attack strategies, techniques, vulnerabilities, and attack patterns, and extracts the relationships between these entities. Based on the identified entities and the extracted relationships, a knowledge graph is constructed.
[0036] Specifically, in this technical solution, the extracted entities include, but are not limited to, multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities. The ATT&CK entities include MitreTactic entities, MitreAttackPattern entities, and MitreMitigations entities. MitreTactic is the attack strategy entity within the ATT&CK model framework, while MitreAttackPattern corresponds to a specific attack technique entity within the ATT&CK model framework, and MitreMitigations entities correspond to mitigation measures for specific attacks.
[0037] Since an attack strategy may include multiple attack techniques, and each attack technique may also have its own sub-attack techniques, the paradigm representation of the knowledge graph entities constructed in step S101 is as follows: Figure 2 The above, Figure 2 This is a schematic diagram illustrating the entity, relationship, and attribute descriptions disclosed in an embodiment of this application. Taking the CWE entity as an example, the entity attributes of the CWE entity include: number (name), Chinese description information, auxiliary description information, and mitigation measures information.
[0038] CWE (Common Weakness Enumeration) refers to a class of vulnerabilities, so for every CWE ID, there are multiple CVE IDs. However, in the CVE dataset, some CVEs lack explicit CWE type labeling due to insufficient information or recent release. Therefore, all CVE entities include a first CVE entity and a second CVE entity. The first CVE entity has a clearly corresponding CWE ID, while the second CVE entity does not.
[0039] Step S2: Input each second CVE entity and the entity attributes corresponding to the second CVE entity into the first text classification model, and output the CWE number corresponding to the second CVE entity.
[0040] In the above steps, the server inputs each second CVE entity and the entity attributes corresponding to the second CVE entity into the first text classification model, and outputs the CWE number corresponding to the second CVE entity.
[0041] Specifically, in this technical solution, the first text classification model is the Fasttext model. The entity attributes corresponding to the second CVE entity include, but are not limited to, the Chinese description information of the vulnerability, the mitigation measures information, and the Chinese description information of the CPE components that it affects.
[0042] Before inputting the entity attributes of each second CVE entity into the first text classification model, it is necessary to perform operations such as word segmentation, stop word removal, and deduplication on the input text in conjunction with the constructed Chinese and English stop word list in the cybersecurity field.
[0043] In the process of obtaining the CWE number corresponding to the second CVE entity, the first text classification model may encounter issues. For some CWEs, the number of CVE entities to which they belong exceeds a first threshold (i.e., too many CVE entities to which some CWEs belong), preferably 100; or for some CWEs, the number of CVE entities to which they belong is less than a second threshold (i.e., too few CVE entities to which some CWEs belong), preferably 20. Therefore, when the number of CVE entities to which some CWEs belong is too high, a random undersampling approach should be used in stratified sampling. When the number of CVE entities to which some CWEs belong is too low, a data augmentation approach should be used in stratified sampling. Data augmentation methods include synonym replacement and word exchange, which can be significantly improved by using a thesaurus constructed by those skilled in the art. For example, if some CWEs belong to too many CVE entities (e.g., 10,000 CVE entities), 100 CVE entities need to be randomly selected from those 10,000. If some CWEs belong to too few CVE entities, an oversampling strategy is needed, such as synonym replacement or randomly changing the word order in the text. Ultimately, the goal is to ensure that the number of CVE entities belonging to each CWE is greater than or equal to the second threshold and less than or equal to the first threshold, and to ensure a uniform distribution of these entities by setting a random seed.
[0044] After the first text classification model is trained on each second CVE entity and the entity corresponding to the second CVE entity using the above method, the first text classification model will output the CWE number corresponding to the second CVE entity.
[0045] Step S3: Based on knowledge graph and graph theory related algorithms, analyze all CVE entities in the CVE entity set and construct the association mapping relationship between all CVE entities and all ATT&CK entities; wherein, the CVE entity set includes multiple first CVE entities and multiple second CVE entities.
[0046] In the above steps, the server analyzes all CVE entities in the CVE entity set based on knowledge graph and graph theory-related algorithms, and constructs the association mapping relationship between all CVE entities and all ATT&CK entities.
[0047] Specifically, Neo4j uses the Cypher query language to analyze the CVE graph. Statistical analysis of CVE entities revealed approximately 155,000 CVE entities with explicit CWE vulnerability type IDs (this number changes daily with the addition and modification of CVE entities; the current number of publicly disclosed CVE entities is approximately 201,000). These 155,000 CVE entities are associated with 376 specific CWE vulnerability type IDs (there are a total of 925 CWE entities in the knowledge graph of this technical solution). Furthermore, based on the mapping relationship between CWE entities and CAPEC entities, the 376 CWE entity types are ultimately directly mapped to 382 specific CAPEC attack techniques (there are a total of 672 CAPEC entities in the knowledge graph of this technical solution). Therefore, in this technical solution, the CVE entity set contains approximately 155,000 first-level CVE entities and 46,000 second-level CVE entities. It should be noted that the number of second-level CVE entities will continue to be updated and grow based on newly discovered system vulnerabilities.
[0048] Reference Figure 2 ,like Figure 2 The CAPEC entities shown are interconnected. Further statistical analysis of the relationships between CWE and CAPEC entities in the knowledge graph, treating these relationships as an undirected graph, reveals that there are no isolated nodes among CAPEC entities; that is, any two CAPEC entities can be connected through a unique path within the graph. Therefore, this provides a theoretical basis for mining the associations of attack techniques between CVE and ATT&CK entities. Subsequent embodiments will detail the specific steps for constructing the mapping relationship between all CVE and ATT&CK entities, so they will not be elaborated upon here.
[0049] Step S4: Output the association mapping relationship.
[0050] In the above steps, the server outputs the association mapping relationship.
[0051] Specifically, in this technical solution, the server outputs the constructed association mapping relationship to the security professional's terminal, which includes, but is not limited to, electronic devices such as computers, mobile phones, and tablets used by the security professional.
[0052] In one possible implementation, the feasibility of constructing an association mapping relationship between all CVE entities and all ATT&CK entities based on knowledge graphs and graph theory-related algorithms is further demonstrated. The knowledge graph also includes multiple CAPEC entities; the association mapping relationship includes multiple first sub-mapping relationships and multiple second sub-mapping relationships. (Refer to...) Figure 3This illustrates a second flowchart of a knowledge graph-based association method for ATT&CK and CVE disclosed in this application. Step S3 specifically includes steps S31-S35, as follows: Step S31: Determine whether the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity.
[0053] Specifically, in this technical solution, the server determines whether the current CWE entity corresponding to each CVE entity maps to at least one CAPEC entity. CAPEC (Common Attack Pattern Enumeration and Classification) is a dataset for enumerating and classifying attack types; it is a classification dataset for commonly used attack types. Therefore, one CWE entity can map to multiple CAPEC entities.
[0054] Step S32: When the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity, determine whether there is a mapping relationship between the current CAPEC entity and the ATT&CK entity.
[0055] Specifically, in this technical solution, when the server determines that the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity, it then determines whether there is a mapping relationship between the current CAPEC entity and the ATT&CK entity.
[0056] Step S33: When there is a mapping relationship between the current CAPEC entity and the ATT&CK entity, the first sub-mapping relationship between the CVE entity and the ATT&CK entity corresponding to the current CAPEC entity is obtained based on the knowledge graph.
[0057] Specifically, in this technical solution, when the server determines that there is a mapping relationship between the current CAPEC entity and the ATT&CK entity, it obtains the first sub-mapping relationship between the CVE entity and the ATT&CK entity corresponding to the current CAPEC entity based on the knowledge graph.
[0058] Step S34: When there is no mapping relationship between the current CAPEC entity and the ATT&CK entity, based on the CAPEC entity set corresponding to the current CWE entity, select the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set.
[0059] Specifically, in this technical solution, when the server determines that there is no mapping relationship between the current CAPEC entity and the ATT&CK entity, it selects the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set corresponding to the current CWE entity.
[0060] Since the current CWE entity maps to at least one CAPEC entity, a CAPEC entity set is constructed based on all CAPEC entities mapped to the current CWE entity. Then, the K first CAPEC entities, ranked by similarity from highest to lowest, are selected from all CAPEC entities. The specific steps for selecting the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set corresponding to the current CWE entity will be described in detail in subsequent embodiments, and therefore will not be elaborated upon here.
[0061] Step S35: Based on the knowledge graph, obtain the second sub-mapping relationship between the CVE entities and the ATT&CK entities corresponding to the K first CAPEC entities.
[0062] Specifically, in this technical solution, the server obtains the second sub-mapping relationship between the CVE entities and ATT&CK entities corresponding to the K first CAPEC entities based on the knowledge graph. The specific method and steps for obtaining the second sub-mapping relationship between the CVE entities and ATT&CK entities corresponding to the K first CAPEC entities based on the knowledge graph will be described in detail in subsequent embodiments, and therefore will not be elaborated here.
[0063] In one possible implementation, refer to Figure 3 Following step S31, the method further includes steps S36-S39, as follows: Step S36: When the current CWE entity corresponding to each CVE entity is not mapped to any CAPEC entity, construct an association set of the first CWE entity associated with the vulnerability type of the current CWE entity.
[0064] Specifically, in this technical solution, when the server determines that the current CWE entity corresponding to each CVE entity does not map to any CAPEC entity, it then searches for other CWE relationships associated with the current CWE entity based on the knowledge graph. Different association mining depths are set for different relationships between CWE entities, such as... Figure 2As shown: Requires indicates that the vulnerability type requires its associated vulnerability, so the relationship between the two CWE entities is very close to a certain extent; ChildOf indicates the relationship of child vulnerabilities, meaning that a certain CWE entity is a child vulnerability of another CWE entity, and their relationship is also very close; CanAlsoBe indicates that the vulnerability type can also be other vulnerability types, etc.; and by limiting the depth of mining and the number of elements in the set, an associated set containing a suitable number of elements is obtained.
[0065] Step S37: When the associated set is a non-empty set, based on the mapping relationship between the first CWE entity and the CAPEC entity, obtain at least one CAPEC entity associated with the current CWE entity.
[0066] Specifically, in this technical solution, when the association set is a non-empty set, it means that the association set contains a first CWE entity associated with the vulnerability type of the current CWE entity. Then, based on the mapping relationship between the first CWE entity and the CAPEC entity, at least one CAPEC entity associated with the current CWE entity is obtained.
[0067] Step S38: When the association set is empty, the first text information is input into the second text classification model to obtain the second CWE entity; the similarity between the first text information corresponding to the second CWE entity and the first text information corresponding to the current CWE entity is greater than the first preset similarity threshold.
[0068] Specifically, in this technical solution, when the association set is empty (i.e., there is no first CWE entity associated with the vulnerability type of the current CWE entity), the first text information is input into the second text classification model to obtain the second CWE entity.
[0069] Since the entity attributes of a CWE entity also include first text information, which consists of the Chinese description, auxiliary description, and mitigation measures information in the entity attributes corresponding to the CWE entity, the second text classification model is also a Fasttext model. To distinguish it from the Fasttext model in the aforementioned embodiments, this Fasttext model is referred to as the second text classification model. After inputting the Chinese description, auxiliary description, and mitigation measures information of the CWE entity into the second text classification model, the second text classification model will use the cosine similarity calculation formula to calculate the text similarity between the first text information corresponding to the current CWE entity and the first text information corresponding to other CWE entities, and filter out CWE entities with a similarity greater than a first preset similarity threshold to obtain the second CWE entity. The first preset similarity threshold can be specifically set according to the actual situation; preferably, the first preset similarity threshold is 70%.
[0070] Step S39: Based on the mapping relationship between the second CWE entity and the CAPEC entity, obtain at least one CAPEC entity associated with the current CWE entity.
[0071] Specifically, in this technical solution, the server obtains at least one CAPEC entity associated with the current CWE entity based on the mapping relationship between the second CWE entity and the CAPEC entity.
[0072] It should be noted that step S32 is executed after step S39 is completed.
[0073] In one possible implementation, refer to Figure 4 This illustrates the third flowchart of a knowledge graph-based association method for ATT&CK and CVE disclosed in this application. Step S34 specifically includes steps S341-S343, as follows: Step S341: Calculate the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity.
[0074] Specifically, in this technical solution, the Jaccard coefficient, also known as the Jaccard similarity coefficient, is used to compare the similarity and differences between a finite set of samples. A higher Jaccard coefficient indicates a higher similarity. Therefore, the server calculates the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity to compare the similarity between the current CAPEC entity and other CAPEC entities.
[0075] In one possible implementation, the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity is calculated using the following formula: Where A is the set of elements in the current CAPEC entity, B is the set of elements in any CAPEC entity in the CAPEC entity set, || is the number of elements in the set, ∩ is the intersection, and ∪ is the union.
[0076] Step S342: Select L second CAPEC entities from the CAPEC entity set; the Jaccard coefficient of the second CAPEC entity and the current CAPEC entity is greater than the preset Jaccard coefficient threshold.
[0077] Specifically, in this technical solution, the server selects L second CAPEC entities from the CAPEC entity set. The value of L is not limited in this application. The preset Jaccard coefficient threshold needs to be set according to the actual situation, and its preferred value is 0.7.
[0078] Step S343: Based on the n-order association relationship corresponding to the second CAPEC entity, filter and converge the L second CAPEC entities to obtain K first CAPEC entities, where n is a positive integer greater than 1.
[0079] Specifically, in this technical solution, since CAPEC entities have direct relationships, the L second CAPEC entities can be further filtered and consolidated based on first-order, second-order, and multi-order relationships to obtain K first CAPEC entities. Here, a first-order relationship refers to other CAPEC entities associated with the current CAPEC entity; a second-order relationship is one associated through an intermediate entity; and a multi-order relationship is one associated through multiple intermediate entities. In a multi-order relationship, the first element is CAPEC, and the intermediate elements can be CWE or CAPEC.
[0080] In one possible implementation, refer to Figure 5 This illustrates the fourth flowchart of a knowledge graph-based association method for ATT&CK and CVE disclosed in this application. Step S35 specifically includes steps S351-S354, as follows: Step S351: Based on the mapping relationship between X directly associated CAPEC entities and ATT&CK entities, obtain the direct sub-mapping relationship.
[0081] Specifically, in this technical solution, the K first CAPEC entities include X directly related CAPEC entities and Y indirectly related CAPEC entities, where X + Y = K; the second sub-mapping relationship includes direct sub-mapping relationships and indirect sub-mapping relationships. That is, X of the K first CAPEC entities are directly related CAPEC entities, which can be directly associated with the ATT&CK entity, thus obtaining a direct sub-mapping relationship.
[0082] Step S352: Input the second text information corresponding to the Y indirectly associated CAPEC entities and the third text information corresponding to the ATT&CK entities into the third text classification model, and calculate the cosine similarity between the second text information and the third text information.
[0083] Specifically, in this technical solution, the second text information consists of the descriptive information and mitigation measures information in the entity attributes corresponding to the CAPEC entity. The third text information consists of the descriptive information and mitigation measures information in the entity attributes corresponding to the ATT&CK entity. The third text classification model is also a Fasttext model. To distinguish it from the two Fasttext models in the aforementioned embodiments, the Fasttext model here is referred to as the third text classification model.
[0084] In one possible implementation, the cosine similarity between the second and third text information is calculated using the following formula: in, This is the embedding vector for the second text information. This is the embedding vector for the third text information. Let C be the magnitude of vector C. Let be the magnitude of vector D.
[0085] Specifically, in this technical solution, vector sum vector The word vectors are calculated by inputting the corresponding second and third text information into a third text classification model. The input second and third text information need to undergo operations such as word segmentation and stop word removal. For word segmentation, jieba segmentation can be used, and the accuracy of segmentation can be improved by using dictionaries built by experts in this field, such as CAPEC and ATT&CK.
[0086] Step S353: Select multiple related ATT&CK entities from all ATT&CK entities. The cosine similarity between the third text information corresponding to each related ATT&CK entity and the second text information corresponding to the indirectly related CAPEC entity is greater than the second preset similarity threshold.
[0087] Specifically, in this technical solution, the server filters out multiple related ATT&CK entities from all ATT&CK entities. The cosine similarity between the third text information corresponding to each related ATT&CK entity and the second text information corresponding to the indirectly related CAPEC entity is greater than a second preset similarity threshold. The second preset similarity threshold needs to be set according to the actual situation, and its preferred value is 0.8.
[0088] Step S354: Based on the mapping relationship between the Y indirectly associated CAPEC entities and the associated ATT&CK entities, obtain the indirect sub-mapping relationship.
[0089] Specifically, in this technical solution, the server obtains indirect sub-mapping relationships based on the mapping relationship between Y indirectly associated CAPEC entities and associated ATT&CK entities.
[0090] To better understand this technical solution, the following embodiments illustrate the construction of the association mapping relationship between CVE entities and ATT&CK entities based on the methods described in the foregoing embodiments: For example, the system vulnerability CVE-2022-2483 has a corresponding CWE entity with the CWE number CWE-1282. This CWE uses the CAPEC attack techniques CAPEC-1 and CAPEC-180. The ATT&CK-related attack technique entries for both CAPEC-1 and CAPEC-180 are T1574.010. Therefore, we can find that the associated ATT&CK attack technique is T1574.010 (i.e., the MitreATT&CKPattern entity within the ATT&CK entity), thus establishing the association between the CVE entity and the ATT&CK entity. Furthermore, we can find the corresponding attack strategy entity (MitreTantic) and mitigation entity (MitreMitigations) in the knowledge graph.
[0091] On the other hand, for example, the system vulnerability CVE-2023-3520 has a corresponding CWE entity with the CWE number CWE-614. This CWE uses the attack technique CAPEC-102, which has no explicit ATT&CK associated entry, and this entity also has no directly associated CAPEC attack technique entity. By searching the knowledge graph using the Jaccard coefficient similarity, the five most similar CAPEC entities are: CAPEC-509 (associated with ATT&CK entity T1558.003), CAPEC-645 (associated with ATT&CK entity T1550.002), CAPEC-644 (associated with ATT&CK entity 1550.003), CAPEC-117 (no explicitly associated ATT&CK entity), and CAPEC-447 (no explicitly associated ATT&CK entity). Using a third text classification model, based on the description and mitigation information corresponding to CAPEC-117 and CAPEC-447, the most similar ATT&CK entities are identified as T1199 and T1195, thus establishing the association between the CVE entity and the ATT&CK entity. Furthermore, after verification by professionals in the field that the attack conforms to technical specifications and requirements, the corresponding attack strategy and mitigation information can be found based on the attack technique number. For example, the attack strategy entities (MitreTantic) corresponding to ATT&CK entity T1550.002 are TA005 and TA008, and the corresponding mitigation entities (MitreMitigations) are M1018, M1026, M1051, and M1052.
[0092] On the other hand, consider the system vulnerability CVE-2022-1734. Its corresponding CWE entity has a CWE ID of CWE-416, and this CWE has no directly associated CAPEC entities. Therefore, by analyzing the correlation between vulnerability types among CWEs, we obtain the most similar association set: {CWE-120, CWE-825, CWE-672}. The CWE entity with CWE ID CWE-672 is associated with multiple CAPEC entities, while the other two CWE entities have no associated CAPEC entities. Based on the multiple CAPEC entities associated with the CWE entity with CWE ID CWE-672, we then search for the mapping relationship between these CAPEC entities and the ATT&CK entity, thus obtaining the association relationship between this CVE entity and the ATT&CK entity.
[0093] This application also provides a knowledge graph-based association system for ATT&CK and CVE, referring to... Figure 6This document illustrates a schematic diagram of a knowledge graph-based association system for ATT&CK and CVE provided in an embodiment of this application. The system includes: a knowledge graph construction module, a processing module, and an output module. The knowledge graph construction module is used to construct a knowledge graph based on the ATT&CK model framework and multiple preset datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities. The multiple CVE entities include first CVE entities and second CVE entities. The entity attributes of a CWE entity include a CWE number, and the first CVE entity has a clearly corresponding CWE number, while the second CVE entity does not have a clearly corresponding CWE number. The processing module is used to input each second CVE entity into a first text classification model and output the CWE number corresponding to the second CVE entity. The processing module is also used to analyze all CVE entities in the CVE entity set based on knowledge graph and graph theory-related algorithms, and construct an association mapping relationship between all CVE entities and all ATT&CK entities. The CVE entity set includes multiple first CVE entities and multiple second CVE entities. The output module is used to output the association mapping relationship.
[0094] In one possible implementation, the processing module is further configured to determine whether the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity; the processing module is further configured to determine whether there is a mapping relationship between the current CAPEC entity and the ATT&CK entity when the current CWE entity corresponding to each CVE entity is mapped to at least one CAPEC entity; the processing module is further configured to obtain a first sub-mapping relationship between the CVE entity and the ATT&CK entity corresponding to the current CAPEC entity based on the knowledge graph when there is a mapping relationship between the current CAPEC entity and the ATT&CK entity; the processing module is further configured to select the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set based on the CAPEC entity set corresponding to the current CWE entity when there is no mapping relationship between the current CAPEC entity and the ATT&CK entity based on the knowledge graph; the processing module is further configured to obtain a second sub-mapping relationship between the CVE entity and the ATT&CK entity corresponding to the K first CAPEC entities based on the knowledge graph.
[0095] In one possible implementation, the processing module is further configured to: construct an association set of first CWE entities associated with the vulnerability type of the current CWE entity when the current CWE entity corresponding to each CVE entity is not mapped to any CAPEC entity; the processing module is further configured to: obtain at least one CAPEC entity associated with the current CWE entity based on the mapping relationship between the first CWE entity and the CAPEC entity when the association set is a non-empty set; the processing module is further configured to: input the first text information into a second text classification model to obtain a second CWE entity when the association set is an empty set; the similarity between the first text information corresponding to the second CWE entity and the first text information corresponding to the current CWE entity is greater than a first preset similarity threshold; and the processing module is further configured to: obtain at least one CAPEC entity associated with the current CWE entity based on the mapping relationship between the second CWE entity and the CAPEC entity.
[0096] In one possible implementation, the processing module is further configured to calculate the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity; the processing module is further configured to select L second CAPEC entities from the CAPEC entity set; the Jaccard coefficient between the second CAPEC entity and the current CAPEC entity is greater than a preset Jaccard coefficient threshold; the processing module is further configured to filter and converge the L second CAPEC entities based on the n-order association relationship corresponding to the second CAPEC entities to obtain K first CAPEC entities, where n is a positive integer greater than 1.
[0097] In one possible implementation, the processing module is further configured to calculate the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity using the following formula: Where A is the set of elements in the current CAPEC entity, B is the set of elements in any CAPEC entity in the CAPEC entity set, || is the number of elements in the set, ∩ is the intersection, and ∪ is the union. In one possible implementation, In one possible implementation, the processing module is further configured to obtain direct sub-mapping relationships based on the mapping relationships between X directly associated CAPEC entities and ATT&CK entities; the processing module is further configured to input the second text information corresponding to Y indirectly associated CAPEC entities and the third text information corresponding to Y indirectly associated CAPEC entities into a third text classification model to calculate the cosine similarity between the second text information and the third text information; the processing module is further configured to filter out multiple associated ATT&CK entities from all ATT&CK entities, wherein the cosine similarity between the third text information corresponding to each associated ATT&CK entity and the second text information corresponding to the indirectly associated CAPEC entity is greater than a second preset similarity threshold; the processing module is further configured to obtain indirect sub-mapping relationships based on the mapping relationships between Y indirectly associated CAPEC entities and associated ATT&CK entities.
[0098] In one possible implementation, the processing module is further configured to calculate the cosine similarity between the second and third text information using the following formula: in, This is the embedding vector for the second text information. This is the embedding vector for the third text information. Let C be the magnitude of vector C. Let be the magnitude of vector D.
[0099] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0100] This application also discloses an electronic device. (See reference...) Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 700 may include: at least one processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.
[0101] The communication bus 702 is used to enable communication between these components.
[0102] The user interface 703 may include a display screen and a camera. Optionally, the user interface 703 may also include a standard wired interface and a wireless interface.
[0103] The network interface 704 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0104] The processor 701 may include one or more processing cores. The processor 701 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 705, and by calling data stored in memory 705. Optionally, the processor 701 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 701 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 701 and may be implemented as a separate chip.
[0105] The memory 705 may include random access memory (RAM) or read-only memory. Optionally, the memory 705 may include a non-transitory computer-readable storage medium. The memory 705 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 705 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 705 may also be at least one storage device located remotely from the aforementioned processor 701. (Refer to...) Figure 7 The memory 705, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program.
[0106] exist Figure 7 In the illustrated electronic device 700, the user interface 703 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 701 can be used to call an application stored in the memory 705. When executed by one or more processors 701, the electronic device 700 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0107] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0108] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0112] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will readily conceive of those skilled in the art upon consideration of the specification and the disclosure of practical truths.
[0113] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A knowledge graph-based association method for ATT&CK and CVE, characterized in that, The method includes: A knowledge graph is constructed based on the ATT&CK model framework and multiple preset datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities. The multiple CVE entities include a first CVE entity and a second CVE entity. The entity attribute of the CWE entity includes a CWE number, and the first CVE entity has a clearly corresponding CWE number, while the second CVE entity does not have a clearly corresponding CWE number. Each second CVE entity and its corresponding entity attribute are input into the first text classification model, and the CWE number corresponding to the second CVE entity is output. Based on the knowledge graph and graph theory related algorithms, all CVE entities in the CVE entity set are analyzed to construct the association mapping relationship between all the CVE entities and all the ATT&CK entities; wherein, the CVE entity set includes multiple first CVE entities and multiple second CVE entities; Output the aforementioned association mapping relationship; The knowledge graph includes multiple CAPEC entities; the association mapping relationship includes multiple first sub-mapping relationships and multiple second sub-mapping relationships; the step of analyzing all CVE entities in the CVE entity set based on the knowledge graph and graph theory-related algorithms, and constructing the association mapping relationship between all CVE entities and all ATT&CK entities, specifically includes: Determine whether the current CWE entity corresponding to each of the CVE entities is mapped to at least one of the CAPEC entities; When each of the CVE entities corresponds to a current CWE entity that is mapped to at least one of the CAPEC entities, determine whether there is a mapping relationship between the current CAPEC entity and the ATT&CK entity. When there is a mapping relationship between the current CAPEC entity and the ATT&CK entity, the first sub-mapping relationship between the CVE entity corresponding to the current CAPEC entity and the ATT&CK entity is obtained based on the knowledge graph. When there is no mapping relationship between the current CAPEC entity and the ATT&CK entity, based on the CAPEC entity set corresponding to the current CWE entity, select the K first CAPEC entities that are ranked first in similarity to the current CAPEC entity from the CAPEC entity set. Based on the knowledge graph, a second sub-mapping relationship is obtained between the CVE entities corresponding to the K first CAPEC entities and the ATT&CK entities.
2. The method according to claim 1, characterized in that, The entity attributes of the CWE entity also include first text information; after determining whether the current CWE entity corresponding to each CVE entity maps to at least one CAPEC entity, the method further includes: When the current CWE entity corresponding to each of the CVE entities is not mapped to any of the CAPEC entities, construct an association set of the first CWE entities associated with the vulnerability type of the current CWE entity; When the association set is a non-empty set, based on the mapping relationship between the first CWE entity and the CAPEC entity, at least one CAPEC entity associated with the current CWE entity is obtained; When the association set is empty, the first text information is input into the second text classification model to obtain the second CWE entity; the similarity between the first text information corresponding to the second CWE entity and the first text information corresponding to the current CWE entity is greater than the first preset similarity threshold. Based on the mapping relationship between the second CWE entity and the CAPEC entity, at least one CAPEC entity associated with the current CWE entity is obtained.
3. The method according to claim 1, characterized in that, The step of selecting the K first CAPEC entities with the highest similarity to the current CAPEC entity from the CAPEC entity set corresponding to the current CWE entity specifically includes: Calculate the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity; From the CAPEC entity set, L second CAPEC entities are selected; the Jaccard coefficient of the second CAPEC entity and the current CAPEC entity is greater than a preset Jaccard coefficient threshold. Based on the n-order association relationship corresponding to the second CAPEC entity, the L second CAPEC entities are filtered and converged to obtain K first CAPEC entities, where n is a positive integer greater than 1.
4. The method according to claim 3, characterized in that, The calculation of the Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity specifically includes: The Jaccard coefficient between any CAPEC entity in the CAPEC entity set and the current CAPEC entity is calculated using the following formula: ; Where A is the set of elements in the current CAPEC entity, and B is the set of elements in any CAPEC entity in the CAPEC entity set. The number of elements in the set. For intersection, It is a union.
5. The method according to claim 1, characterized in that, The K first CAPEC entities include X directly related CAPEC entities and Y indirectly related CAPEC entities, where X + Y = K; the second sub-mapping relationship includes direct sub-mapping relationships and indirect sub-mapping relationships; the step of obtaining the second sub-mapping relationship between the CVE entities corresponding to the K first CAPEC entities and the ATT&CK entities based on the knowledge graph specifically includes: Based on the mapping relationship between X directly associated CAPEC entities and the ATT&CK entity, the direct sub-mapping relationship is obtained; Input the second text information corresponding to the Y indirectly associated CAPEC entities and the third text information corresponding to the ATT&CK entities into the third text classification model, and calculate the cosine similarity between the second text information and the third text information; Multiple associated ATT&CK entities are selected from all the ATT&CK entities, and the cosine similarity between the third text information corresponding to each associated ATT&CK entity and the second text information corresponding to the indirectly associated CAPEC entity is greater than a second preset similarity threshold. Based on the mapping relationship between the Y indirectly associated CAPEC entities and the associated ATT&CK entities, the indirect sub-mapping relationship is obtained.
6. The method according to claim 5, characterized in that, The calculation of the cosine similarity between the second text information and the third text information specifically includes: The cosine similarity between the second and third text information is calculated using the following formula: ,in, This is the embedding vector of the second text information. The embedding vector of the third text information. Let C be the magnitude of vector C. Let be the magnitude of vector D.
7. A knowledge graph-based association system for ATT&CK and CVE, characterized in that, The system for executing the knowledge graph-based association method for ATT&CK and CVE as described in claim 1 includes: a knowledge graph construction module, a processing module, and an output module; The knowledge graph construction module is used to construct a knowledge graph based on the ATT&CK model framework and multiple preset datasets. The knowledge graph includes multiple CVE entities, multiple CWE entities, and multiple ATT&CK entities. The multiple CVE entities include a first CVE entity and a second CVE entity. The entity attribute of the CWE entity includes a CWE number, and the first CVE entity has a clearly corresponding CWE number, while the second CVE entity does not have a clearly corresponding CWE number. The processing module is used to input each of the second CVE entities into the first text classification model and output the CWE number corresponding to the second CVE entity. The processing module is further configured to analyze all CVE entities in the CVE entity set based on the knowledge graph and graph theory-related algorithms, and construct an association mapping relationship between all CVE entities and all ATT&CK entities; wherein, the CVE entity set includes multiple first CVE entities and multiple second CVE entities; The output module is used to output the association mapping relationship.
8. An electronic device, characterized in that, The device includes a processor (701), a memory (705), a user interface (703), and a network interface (704). The memory (705) is used to store instructions. The user interface (703) and the network interface (704) are used to communicate with other devices. The processor (701) is used to execute the instructions stored in the memory (705) to cause the electronic device (700) to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-source heterogeneous network security knowledge graph construction method and device
CN112131882A
Network security defense scheme generation method based on knowledge graph
CN115510179A