A high-risk vulnerability component association word mining method, device, equipment and medium

CN117435643BActive Publication Date: 2026-09-11GLOBAL ENERGY INTERCONNECTION RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311465365.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2026-09-11
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

[0003]有鉴于此,本发明提供了一种高危漏洞组件关联词挖掘方法、装置、设备及介质,以解决筛选高危害等级漏洞组件关联词需要花费大量的时间和成本的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117435643B_ABST
    Figure CN117435643B_ABST
Patent Text Reader

Abstract

This invention relates to the field of network security technology and discloses a method, apparatus, device, and medium for mining high-risk vulnerability component association words. The method includes: acquiring a dataset containing information corresponding to components and vulnerabilities; constructing a knowledge graph based on the dataset and a preset knowledge graph paradigm; and mining high-risk vulnerability association words for any component using a vulnerability severity rating standard and the TF-IDF algorithm, based on the dataset, knowledge graph, and a preset multidimensional corpus. The preset multidimensional corpus consists of multiple corpora constructed according to various classification methods of components. This invention achieves the sorting and integration of large amounts of relevant data, utilizes the TF-IDF algorithm to mine component association words, and combines a knowledge graph and a vulnerability severity rating standard to automatically mine high-risk vulnerability component association words. It solves the problem in related technologies where screening high-risk vulnerability component association words requires significant time and cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, specifically to a method, apparatus, device, and medium for mining related terms of high-risk vulnerability components. Background Technology

[0002] With the rapid development of the software industry, open-source components are widely used, especially database and operating system components, which are extensively applied in business systems. Component security has become increasingly important. Components are the primary agents through which vulnerabilities ultimately affect systems. In practice, component names are determined by vendors according to their own specifications; aside from the basic Common Platform Enumeration (CPE) specification, there are no other unified naming conventions. Currently, the main method for searching related information based on component names is to perform full-word matching across multiple datasets. However, this method results in too much vulnerability information, requiring significant time and resources to filter out information related to high-risk vulnerability components. Therefore, existing technologies suffer from the problem of consuming substantial time and resources to filter keywords associated with high-risk vulnerability components. Summary of the Invention

[0003] In view of this, the present invention provides a method, apparatus, device and medium for mining high-risk vulnerability component related terms, so as to solve the problem that screening high-risk vulnerability component related terms requires a lot of time and cost.

[0004] In a first aspect, the present invention provides a method for mining high-risk vulnerability component association words, comprising: acquiring a dataset, wherein the dataset contains information corresponding to components and vulnerabilities respectively; constructing a knowledge graph based on the dataset and a preset knowledge graph paradigm, wherein the preset knowledge graph paradigm is used to represent entities, attributes and relationships in the knowledge graph; and mining high-risk vulnerability association words of any component based on the dataset, the knowledge graph and a preset multidimensional corpus using a vulnerability severity rating standard and the TF-IDF algorithm, wherein the preset multidimensional corpus consists of multiple corpora constructed according to various classification methods of components.

[0005] In this embodiment of the invention, by acquiring a dataset, a knowledge graph is constructed based on a preset knowledge graph paradigm for components and vulnerability information. This enables the sorting and integration of a large amount of related data. The TF-IDF algorithm is used to mine the associated words of components, and the knowledge graph and vulnerability severity rating criteria are combined to automatically mine the associated words of high-risk vulnerability components. This solves the problem in related technologies where screening associated words of high-severity vulnerability components requires a significant amount of time and cost.

[0006] In one optional implementation, the dataset includes: vulnerability articles. Before mining high-risk vulnerability-related terms for any component using the vulnerability severity rating standard and the TF-IDF algorithm based on the dataset, knowledge graph, and pre-set multidimensional corpus, the method further includes: extracting vulnerability information from multiple security dimensions based on the vulnerability articles; determining the vulnerability's index values ​​in multiple security dimensions based on the vulnerability's information in multiple security dimensions; and storing the vulnerability's information in multiple security dimensions and the vulnerability's index values ​​in multiple security dimensions in the knowledge graph.

[0007] In this embodiment of the invention, useful information about component vulnerabilities in multiple security dimensions is mined from vulnerability articles, which serves as the basis for screening high-risk vulnerability related terms for components, thereby improving the efficiency of component related term mining; information about vulnerabilities in multiple security dimensions and the indicator values ​​of vulnerabilities in multiple security dimensions are stored in a knowledge graph, realizing the integration and effective utilization of information, further improving mining efficiency.

[0008] In one optional implementation, the knowledge graph includes vulnerability index values ​​across multiple security dimensions. High-risk vulnerability-related terms for any component are mined using a vulnerability severity rating standard and the TF-IDF algorithm, based on the dataset, knowledge graph, and a pre-defined multi-dimensional corpus. This includes: obtaining the dataset corresponding to any component and acquiring the vulnerability index values ​​across multiple security dimensions for that component based on the knowledge graph; calculating the TF-IDF values ​​corresponding to multiple word segments in the dataset corresponding to any component based on the pre-defined multi-dimensional corpus, where word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component; using the weighted sum of the TF-IDF value corresponding to any word segment and the corresponding vulnerability index values ​​across multiple security dimensions as the corresponding high-risk vulnerability score; and determining the high-risk vulnerability-related terms for any component based on the vulnerability severity rating standard and the high-risk vulnerability score.

[0009] In this embodiment of the invention, after mining the associated words of any component from a multidimensional corpus using the TF-IDF algorithm, the high-risk vulnerability associated words of the component are determined by combining the indicator values ​​of the vulnerability in multiple security dimensions in the knowledge graph. This achieves the goal of obtaining the associated words of the component at the high-risk vulnerability level from a large amount of data, and improves the efficiency and accuracy of mining.

[0010] In one optional implementation, the dataset further includes: vulnerability description information and component description information. The TF-IDF values ​​of multiple word segments in the dataset corresponding to any component are calculated based on a preset multidimensional corpus. Word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component. This includes: obtaining vulnerability description information and component description information corresponding to any component; performing stop word removal and word segmentation on the vulnerability description information and component description information to obtain multiple word segments; and calculating the TF-IDF value of any word segment in the corresponding preset multidimensional corpus to obtain the TF-IDF values ​​of multiple word segments in the preset multidimensional corpus.

[0011] In this embodiment of the invention, the TF-IDF value is used to represent the vulnerability description information corresponding to the component and the importance of word segmentation in the component description information, thereby realizing the extraction of useful information of the component.

[0012] In one optional implementation, the weighted sum of the TF-IDF value corresponding to any word segment and the corresponding vulnerability's index values ​​across multiple security dimensions is used as the corresponding high-risk vulnerability score. This includes: selecting a preset number of words with the highest TF-IDF values ​​from a preset multidimensional corpus as target words; and using the weighted sum of the TF-IDF value corresponding to the target word and the corresponding vulnerability's index values ​​across multiple security dimensions as the corresponding high-risk vulnerability score.

[0013] In this embodiment of the invention, since the choice of corpus has a significant impact on the TF-IDF value, different corpora are selected from multiple dimensions to calculate the TF-IDF value, thereby achieving multi-dimensional and fine-grained component association word mining and improving the accuracy of the mining results.

[0014] In one optional implementation, the vulnerability severity rating standard is determined based on a general vulnerability rating system, and when the score of a high-risk vulnerability is greater than a preset value, it is determined to be a high-risk vulnerability.

[0015] In this embodiment of the invention, the general vulnerability scoring system is determined by the official authorities based on various information, such as vulnerability attack methods, attack complexity, and the permissions required for the attack. The vulnerability severity rating standard determined by this system has higher reliability and universality, and the high-risk vulnerability component association words discovered using this rating standard are also more accurate.

[0016] Secondly, the present invention provides a high-risk vulnerability component association word mining device, comprising: a dataset acquisition module for acquiring a dataset, wherein the dataset contains information corresponding to components and vulnerabilities respectively; a knowledge graph construction module for constructing a knowledge graph based on the dataset and a preset knowledge graph paradigm, wherein the preset knowledge graph paradigm is used to represent entities, attributes and relationships in the knowledge graph; and an association word mining module for mining high-risk vulnerability association words of any component based on the dataset, the knowledge graph and a preset multidimensional corpus using a vulnerability severity rating standard and the TF-IDF algorithm, wherein the preset multidimensional corpus consists of multiple corpora constructed according to various classification methods of components.

[0017] In one optional implementation, the dataset includes: vulnerability articles, and the device further includes: a security information extraction module for extracting vulnerability information in multiple security dimensions based on the vulnerability articles; an indicator value determination module for determining the indicator values ​​of the vulnerability in multiple security dimensions based on the vulnerability information in multiple security dimensions; and an indicator value storage module for storing the vulnerability information in multiple security dimensions and the vulnerability indicator values ​​in multiple security dimensions in a knowledge graph.

[0018] In one optional implementation, the knowledge graph includes indicator values ​​of vulnerabilities across multiple security dimensions. The associated word mining module includes: an indicator value acquisition unit, used to acquire the dataset corresponding to any component and acquire the indicator values ​​of vulnerabilities corresponding to any component across multiple security dimensions based on the knowledge graph; a TF-IDF value calculation unit, used to calculate the TF-IDF values ​​corresponding to multiple words in the dataset corresponding to any component based on a preset multi-dimensional corpus, where word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component; a high-risk vulnerability score calculation unit, used to calculate the weighted sum of the TF-IDF value corresponding to any word and the indicator values ​​of the corresponding vulnerability across multiple security dimensions as the corresponding high-risk vulnerability score; and a high-risk vulnerability associated word determination unit, used to determine the high-risk vulnerability associated words of any component based on the vulnerability severity level scoring standard and the high-risk vulnerability score.

[0019] In one optional implementation, the dataset further includes: vulnerability description information and component description information. The TF-IDF value calculation unit includes: an information acquisition subunit, used to acquire vulnerability description information and component description information corresponding to any component; a preprocessing subunit, used to perform stop word removal and word segmentation processing on the vulnerability description information and component description information to obtain multiple word segments; and a TF-IDF value calculation subunit, used to calculate the TF-IDF value of any word segment in the corresponding preset multidimensional corpus to obtain the TF-IDF values ​​of multiple word segments in the preset multidimensional corpus.

[0020] In one optional implementation, the high-risk vulnerability scoring unit includes: a filtering subunit, used to filter out the preset words with the highest TF-IDF values ​​in the preset multidimensional corpus as target words; and a high-risk vulnerability scoring unit, used to use the weighted sum of the TF-IDF value corresponding to the target word and the indicator values ​​of the corresponding vulnerability in multiple security dimensions as the corresponding high-risk vulnerability scoring value.

[0021] In one optional implementation, the vulnerability severity rating standard is determined based on a general vulnerability rating system, and when the score of a high-risk vulnerability is greater than a preset value, it is determined to be a high-risk vulnerability.

[0022] Thirdly, the present invention provides a computer device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the high-risk vulnerability component association word mining method described in the first aspect or any corresponding embodiment thereof.

[0023] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the high-risk vulnerability component association word mining method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0024] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating the method for mining related terms of high-risk vulnerability components according to an embodiment of the present invention;

[0026] Figure 2 This is a flowchart illustrating another high-risk vulnerability component associated word mining method according to an embodiment of the present invention;

[0027] Figure 3 It is a preset knowledge graph paradigm diagram according to an embodiment of the present invention;

[0028] Figure 4 This is an overall framework diagram of the high-risk vulnerability component associated word mining method according to an embodiment of the present invention;

[0029] Figure 5 This is a structural block diagram of a high-risk vulnerability component related word mining device according to an embodiment of the present invention;

[0030] Figure 6 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] It should be noted that in the description of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0033] According to an embodiment of the present invention, a method for mining related terms of high-risk vulnerability components is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0034] This embodiment provides a method for mining related terms of high-risk vulnerability components, which can be used in the aforementioned mobile terminals, such as central processing units and servers. Figure 1 This is a flowchart illustrating the high-risk vulnerability component association word mining method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0035] Step S101: Obtain the dataset, which contains information corresponding to the components and vulnerabilities respectively. Optionally, the dataset can be a publicly available dataset of components and vulnerabilities, such as: Common Vulnerabilities Exposures (CVE), Common Platform Enumeration (CPE), Common Weakness Enumeration (CWE); or it can be relevant information about components and vulnerabilities obtained from the Internet, such as related articles.

[0036] Step S102: Construct a knowledge graph based on the dataset and a preset knowledge graph paradigm. The preset knowledge graph paradigm is used to represent entities, attributes, and relationships in the knowledge graph. Optionally, the preset knowledge graph paradigm defines the entities, attributes, and relationships in the knowledge graph. The dataset obtained in step S101 is used to construct a corresponding knowledge graph according to the preset knowledge graph paradigm to integrate the data and facilitate searching.

[0037] Step S103: Using the vulnerability severity rating standard and the TF-IDF algorithm, high-risk vulnerability-related terms for any component are mined based on the dataset, knowledge graph, and a pre-set multidimensional corpus. The pre-set multidimensional corpus consists of multiple corpora constructed according to various component classification methods. Optionally, the pre-set multidimensional corpus is a multidimensional corpus constructed from multiple levels, such as multiple corpora constructed from the perspectives of component category, vulnerability type, and vendor type. When mining related terms for any component, the TF-IDF algorithm is first used to find related terms for any component in the pre-set multidimensional corpus. Then, the vulnerability severity rating standard is combined with information from the knowledge graph to mine related terms for components containing high-risk vulnerabilities.

[0038] In this embodiment of the invention, by acquiring a dataset, a knowledge graph is constructed based on a preset knowledge graph paradigm for components and vulnerability information. This enables the sorting and integration of a large amount of related data. The TF-IDF algorithm is used to mine the associated words of components, and the knowledge graph and vulnerability severity rating criteria are combined to automatically mine the associated words of high-risk vulnerability components. This solves the problem in related technologies where screening associated words of high-severity vulnerability components requires a significant amount of time and cost.

[0039] This embodiment provides a method for mining related terms of high-risk vulnerability components, which can be used in the aforementioned mobile terminals, such as central processing units and servers. Figure 2 This is a flowchart of another high-risk vulnerability component associated word mining method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0040] Step S201: Obtain the dataset, which contains information corresponding to components and vulnerabilities respectively. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0041] Step S202: Construct a knowledge graph based on the dataset and a preset knowledge graph paradigm. The preset knowledge graph paradigm is used to represent entities, attributes, and relationships in the knowledge graph. Optionally, the preset knowledge graph paradigm is as follows: Figure 3 As shown in Table 1:

[0042] Table 1. Entities and their corresponding attribute descriptions

[0043]

[0044] The entities CWE, vuln_article, CVE, AffectingComponents, and CVE_references represent vulnerability type, vulnerability article, vulnerability, affected components, and vulnerability information, respectively. Each type of entity has its own attributes. For example, the attributes of the vulnerability entity CVE include: name, Chinese description of the vulnerability, the impact vector of the vulnerability's exploitability indicators in the Common Vulnerability Scoring System (CVSS) standard, and the vulnerability's score. The impact vector of the vulnerability's exploitability indicators in the CVSS standard includes, for example, attack method, attack complexity, and whether UI interaction is required. The vulnerability score is determined according to the CVSS score. The relationships between various entity types in Table 1 include: IntelligenceOf, Mapping, VulnOf, and UsedBy, representing intelligence relationships, mapping relationships, belonging relationships, and used relationships, respectively. Figure 3 As shown, vuln_article-IntelligenceOf-CVE indicates that vuln_article (vulnerable article entity) is a CVE (vulnerable entity) intelligence, and CVE-VulnOf-AffectingComponents indicates that the CVE (vulnerable entity) belongs to an AffectingComponents (affecting component entity) vulnerability. Based on this preset knowledge graph paradigm and the obtained dataset, a corresponding knowledge graph can be constructed.

[0045] It should be noted that the extraction of entities, attributes, and relationships in the knowledge graph can be performed using methods such as machine learning. In some optional implementations, before step S203, the method further includes: extracting vulnerability information from multiple security dimensions based on vulnerability articles; determining the vulnerability's indicator values ​​across multiple security dimensions based on this information; and storing the vulnerability information and indicator values ​​in the knowledge graph. Specifically, after obtaining the vulnerability article, useful information is extracted from it from multiple security dimensions (such as whether there are mitigation measures, patches, exploits, and the readability of the vulnerability analysis), and the indicator values ​​are determined based on a general vulnerability scoring system. This information and indicator values ​​are then stored in the knowledge graph, which can use Neo4j as the storage database.

[0046] Step S203 involves using a vulnerability severity rating standard and the TF-IDF algorithm to mine high-risk vulnerability-related words for any component based on the dataset, knowledge graph, and a pre-set multidimensional corpus. The pre-set multidimensional corpus consists of multiple corpora constructed according to various classification methods of components. Specifically, step S203 includes:

[0047] Step S2031: Obtain the dataset corresponding to any component and obtain the indicator values ​​of the vulnerability corresponding to any component in multiple security dimensions based on the knowledge graph. Optionally, based on the knowledge graph constructed in step S202, the associated vulnerabilities of any component can be determined, and the stored information on vulnerabilities in multiple security dimensions and the indicator values ​​of vulnerabilities in multiple security dimensions can be obtained.

[0048] Step S2032: Calculate the TF-IDF values ​​of multiple word segments in the dataset corresponding to any component based on a pre-defined multidimensional corpus. Word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component. The pre-defined multidimensional corpus may be constructed from three dimensions: component category (software, hardware, operating system), vulnerability category (CWE category to which it belongs), and vendor name. Each dimension constitutes a complete corpus. In some optional implementations, step S2032 includes:

[0049] Step a1: Obtain the vulnerability description information and component description information corresponding to any component.

[0050] Step a2 involves removing stop words and segmenting the vulnerability description information and component description information to obtain multiple segmented words.

[0051] Step a3: Calculate the TF-IDF value of any word segment in the corresponding preset multidimensional corpus, and obtain the TF-IDF values ​​of multiple word segments in the preset multidimensional corpus.

[0052] When calculating TF-IDF values ​​using the TF-IDF algorithm, the entire corpus needs to be considered holistically. Components of the same vulnerability type, component type, and manufacturer should have similar information expressions and similar related words. Therefore, step S2032 calculates the TF-IDF values ​​of multiple word segments obtained after stop word removal and word segmentation in their corresponding similar corpora to obtain similar related words. In this embodiment of the invention, since the choice of corpus has a significant impact on the TF-IDF value, different corpora are selected from multiple dimensions to calculate the TF-IDF value, thereby achieving multi-dimensional and fine-grained component related word mining and improving the accuracy of the mining results.

[0053] Step S2033 involves taking the weighted sum of the TF-IDF value corresponding to any word segment and the corresponding vulnerability's index values ​​across multiple security dimensions as the corresponding high-risk vulnerability score. In some optional implementations, step S2033 includes:

[0054] Step b1 involves selecting the pre-defined word segments with the highest TF-IDF values ​​from the pre-defined multidimensional corpus as target word segments. Optionally, based on the component-related information to be mined, the pre-defined word segments with the highest TF-IDF values ​​from the corresponding pre-defined multidimensional corpus are selected as target word segments, i.e., the component-related words initially mined.

[0055] Step b2 involves using the weighted sum of the TF-IDF value corresponding to the target word and the corresponding vulnerability's indicator values ​​across multiple security dimensions as the high-risk vulnerability score. Optionally, the indicator values ​​of the vulnerability corresponding to the target word across multiple security dimensions can be searched in the knowledge graph. Then, the TF-IDF value of any target word and its corresponding vulnerability's indicator values ​​across multiple security dimensions are weighted and summed, and the result of the weighted sum is used as the corresponding high-risk vulnerability score.

[0056] Step S2034: Determine high-risk vulnerability related terms for any component based on the vulnerability severity rating standard and the high-risk vulnerability score value. Optionally, the vulnerability severity rating standard is determined according to a general vulnerability scoring system. When the high-risk vulnerability score value is greater than a preset value, it is determined to be a high-risk vulnerability related term. Specifically, the knowledge graph contains the vulnerability scores (CVSS attribute values) determined by the general vulnerability scoring system. In this embodiment, a corresponding vulnerability severity rating standard is formulated based on the vulnerability scoring system. As an example of a rating, CVSS is the severity rating of the vulnerability entity CVE, ranging from 0.0 to 10.0. The higher the score, the higher the severity of the vulnerability and the greater the impact. In this embodiment, vulnerabilities with a CVSS of 7.0 or higher are identified as high-risk vulnerabilities. Similarly, a threshold can be set for the high-risk vulnerability score value. If the high-risk vulnerability score value is greater than the threshold, relevant information (such as relevant target words and relevant entities in the knowledge graph) is used as the mined component high-risk vulnerability related terms.

[0057] In some alternative implementations, Figure 4 This is an overall framework diagram of the high-risk vulnerability component association word mining method according to an embodiment of the present invention. In this embodiment, a knowledge graph is first constructed based on datasets such as CVE, CPE, and CWE; then, the TF-IDF language model is used to perform semantic parsing on relevant texts such as component descriptions, vulnerability descriptions, and vulnerability analysis articles; finally, the constructed knowledge graph, combined with the calculated TF-IDF value, is used to mine association words of high-risk vulnerability-related components according to preset expert experience rules (such as constructing vulnerability severity level scoring standards).

[0058] Specifically, the CVE dataset mentioned above is used to standardize the identification of known network threats. It is an international standard for expressing vulnerabilities. By extracting entity information related to vulnerability entities in the knowledge graph, we can explore their associations with entities such as CVEs, vulnerability articles, and components. The CPE dataset is an international standard for naming components, providing the name, specific version, platform, and other information of the software affected by the vulnerability. For example, in "cpe:2.3:a:apache:oozie:5.2.1:*:*:*:*:*:*:*, the CPE standard version is 2.3; 'a' represents the component type as application, i.e., software type. Similarly, 'o' represents the operating system, 'h' represents the hardware; 'apache' indicates the component manufacturer; 'oozie' represents the component name, and '5.2.1' represents the affected component version. The AffectingComponents entity can be constructed using the CPE dataset. The description of this component is as follows: Oozie is an open-source workflow and collaboration service engine based on Apache Hadoop data processing tasks. Oozie is a scalable and extensible data-oriented service running on the Hadoop platform. Oozie includes an offline Hadoop workflow solution and a query processing API. The CWE dataset identifies the specific vulnerability type of a CVE, such as CWE-20 indicating improper input validation, CWE-79 indicating improper input neutralization, and CWE-89 indicating an SQL injection vulnerability, etc. The vuln_article entity consists of vulnerability-related analysis articles obtained from the internet; through entity identification and extraction, its correspondence with CVEs can be determined.

[0059] like Figure 4As shown, the input component name is used to mine component-related terms using the TF-IDF algorithm. For a corpus containing multiple documents, such as a pre-set multi-dimensional corpus built according to vulnerability type, component category, and component vendor, with multiple corpora for each dimension, the TF-IDF value of the component-related terms in each corpus is calculated from multiple dimensions. This allows for fine-grained and precise mining of component-related terms. TF-IDF is the probability that any word appears in all documents. The TF-IDF value of word 'a' can be calculated using the following formula:

[0060] TF-IDF a =TF a *IDF a

[0061]

[0062]

[0063] The higher the TF-IDF value of a word, the higher its frequency of occurrence in the current document, and the less data in the entire corpus contains data related to this document. In other words, words with higher TF-IDF values ​​are more important to the current document and better express its information. As can be seen from the formula above, the choice of corpus has a significant impact on the TF-IDF value. Therefore, this embodiment selects different corpora from multiple dimensions to calculate TF-IDF values, achieving multi-dimensional and fine-grained word mining. Before calculating the TF-IDF value, the corpus and component-related information are preprocessed. For example, based on a self-built stop word list, stop words are removed from vulnerability descriptions and component descriptions, and word segmentation is performed. Then, the TF-IDF values ​​for different corpora are calculated separately. For details, please refer to [link to relevant documentation]. Figure 2 Step S2032 of the illustrated embodiment will not be described again here. TF-IDF is converged at the component description and vulnerability layer description information level, for example, by selecting the multiple words with the highest calculated TF-IDF values.

[0064] On the other hand, knowledge graphs are used to determine whether there are mitigation measures, patches, and exploits for corresponding vulnerabilities based on the input component names, and to perform vulnerability analysis readability assessments. Specifically, corresponding article data (including article content, readability, and other indicators) can be extracted and stored in the knowledge graph, and related information can be selected by combining the highest TF-IDF values ​​selected in the above steps. Optionally, the information mined from vulnerability articles can also be scored. For example, high-risk vulnerabilities with a CVSS score of 7.0 or higher can be screened and statistically analyzed, and the statistically analyzed relevant information can be weighted and statistically analyzed with the TF-IDF values ​​mined from the multi-dimensional corpus to obtain the final high-risk vulnerability score.

[0065] This embodiment utilizes knowledge graphs and word vector models to mine related terms for components from different dimensions. It also combines information from vulnerability articles across multiple security dimensions, such as remediation solutions, vulnerability exploitation, and vulnerability analysis, to score the severity level of vulnerabilities. This allows for effective preference for high-risk vulnerabilities, obtaining a comprehensive description of the component at the high-risk vulnerability level. This improves the comprehensiveness and efficiency of the mining process, solving the problem in related technologies where filtering related terms for high-risk vulnerability components requires significant time and cost.

[0066] This embodiment also provides a high-risk vulnerability component related word mining device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0067] This embodiment provides a device for mining related terms of high-risk vulnerability components, such as... Figure 5 As shown, it includes: a dataset acquisition module 501, used to acquire a dataset, which contains information corresponding to components and vulnerabilities respectively; a knowledge graph construction module 502, used to construct a knowledge graph based on the dataset and a preset knowledge graph paradigm, the preset knowledge graph paradigm being used to represent entities, attributes, and relationships in the knowledge graph; and a related word mining module 503, used to mine high-risk vulnerability related words for any component based on the dataset, knowledge graph, and preset multidimensional corpus using vulnerability severity level scoring standards and the TF-IDF algorithm, the preset multidimensional corpus being multiple corpora constructed according to various classification methods of components.

[0068] In one optional implementation, the dataset includes: vulnerability articles, and the device further includes: a security information extraction module for extracting vulnerability information in multiple security dimensions based on the vulnerability articles; an indicator value determination module for determining the indicator values ​​of the vulnerability in multiple security dimensions based on the vulnerability information in multiple security dimensions; and an indicator value storage module for storing the vulnerability information in multiple security dimensions and the vulnerability indicator values ​​in multiple security dimensions in a knowledge graph.

[0069] In one optional implementation, the knowledge graph includes indicator values ​​of vulnerabilities across multiple security dimensions. The associated word mining module includes: an indicator value acquisition unit, used to acquire the dataset corresponding to any component and acquire the indicator values ​​of vulnerabilities corresponding to any component across multiple security dimensions based on the knowledge graph; a TF-IDF value calculation unit, used to calculate the TF-IDF values ​​corresponding to multiple words in the dataset corresponding to any component based on a preset multi-dimensional corpus, where word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component; a high-risk vulnerability score calculation unit, used to calculate the weighted sum of the TF-IDF value corresponding to any word and the indicator values ​​of the corresponding vulnerability across multiple security dimensions as the corresponding high-risk vulnerability score; and a high-risk vulnerability associated word determination unit, used to determine the high-risk vulnerability associated words of any component based on the vulnerability severity level scoring standard and the high-risk vulnerability score.

[0070] In one optional implementation, the dataset further includes: vulnerability description information and component description information. The TF-IDF value calculation unit includes: an information acquisition subunit, used to acquire vulnerability description information and component description information corresponding to any component; a preprocessing subunit, used to perform stop word removal and word segmentation processing on the vulnerability description information and component description information to obtain multiple word segments; and a TF-IDF value calculation subunit, used to calculate the TF-IDF value of any word segment in the corresponding preset multidimensional corpus to obtain the TF-IDF values ​​of multiple word segments in the preset multidimensional corpus.

[0071] In one optional implementation, the high-risk vulnerability scoring unit includes: a filtering subunit, used to filter out the preset words with the highest TF-IDF values ​​in the preset multidimensional corpus as target words; and a high-risk vulnerability scoring unit, used to use the weighted sum of the TF-IDF value corresponding to the target word and the indicator values ​​of the corresponding vulnerability in multiple security dimensions as the corresponding high-risk vulnerability scoring value.

[0072] In one optional implementation, the vulnerability severity rating standard is determined based on a general vulnerability rating system, and when the score of a high-risk vulnerability is greater than a preset value, it is determined to be a high-risk vulnerability.

[0073] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0074] In this embodiment, the high-risk vulnerability component related word mining device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0075] This invention also provides a computer device having the above-described features. Figure 5 The high-risk vulnerability component associated word mining device shown.

[0076] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 6 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.

[0077] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0078] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0079] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0080] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0081] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0082] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0083] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for mining related terms of high-risk vulnerability components, characterized in that, The method includes: Obtain the dataset, which contains information corresponding to components and vulnerabilities respectively; A knowledge graph is constructed based on the dataset and a preset knowledge graph paradigm, wherein the preset knowledge graph paradigm is used to represent entities, attributes, and relationships in the knowledge graph; wherein the knowledge graph includes indicator values ​​of vulnerabilities in multiple security dimensions; Using a vulnerability severity rating standard and the TF-IDF algorithm, high-risk vulnerability-related words for any component are mined from the dataset, the knowledge graph, and a pre-set multidimensional corpus. The pre-set multidimensional corpus comprises multiple corpora constructed according to various classification methods of components. The process of mining high-risk vulnerability-related words for any component using the vulnerability severity rating standard and the TF-IDF algorithm includes: Obtain the dataset corresponding to any component and, based on the knowledge graph, obtain the indicator values ​​of the vulnerability corresponding to any component in multiple security dimensions; The TF-IDF values ​​of multiple word segments in the dataset corresponding to any component are calculated based on a preset multidimensional corpus. The word segments are multiple words obtained after preprocessing the dataset corresponding to any component. The weighted sum of the TF-IDF value corresponding to any word and the corresponding vulnerability's index values ​​across multiple security dimensions is used as the corresponding high-risk vulnerability score. Based on the vulnerability severity rating criteria and the high-risk vulnerability rating, determine the associated keywords for any component's high-risk vulnerabilities.

2. The method for mining related terms of high-risk vulnerability components according to claim 1, characterized in that, The dataset includes: vulnerability articles. Before mining high-risk vulnerability-related words for any component using the vulnerability severity rating standard and the TF-IDF algorithm based on the dataset, the knowledge graph, and the preset multidimensional corpus, the method further includes: Information about the vulnerability across multiple security dimensions was extracted from the vulnerability article. Based on the information about the vulnerability across multiple security dimensions, determine the indicator values ​​of the vulnerability in multiple security dimensions; Information about the vulnerability across multiple security dimensions, as well as the indicator values ​​of the vulnerability across multiple security dimensions, are stored in a knowledge graph.

3. The method for mining related terms of high-risk vulnerability components according to claim 1, characterized in that, The dataset also includes: vulnerability description information and component description information. The step of calculating the TF-IDF values ​​of multiple word segments in the dataset corresponding to any component based on a preset multidimensional corpus, wherein the word segmentation refers to multiple words obtained after preprocessing the dataset corresponding to any component, including: Obtain the vulnerability description information and component description information corresponding to any component; The vulnerability description information and component description information are processed by removing stop words and segmenting words to obtain multiple word segments; Calculate the TF-IDF value of any word segment in the corresponding preset multidimensional corpus, and obtain the TF-IDF values ​​of multiple word segments in the preset multidimensional corpus.

4. The method for mining related terms of high-risk vulnerability components according to claim 3, characterized in that, The method of using the weighted sum of the TF-IDF value corresponding to any word segment and the corresponding vulnerability's index values ​​across multiple security dimensions as the corresponding high-risk vulnerability score includes: Each of the preset word segments with the highest TF-IDF values ​​in the preset multidimensional corpus is selected as the target word segment; The weighted sum of the TF-IDF value corresponding to the target word and the corresponding vulnerability's index values ​​across multiple security dimensions is used as the corresponding high-risk vulnerability score.

5. The method for mining related terms of high-risk vulnerability components according to claim 1, characterized in that, The vulnerability severity rating criteria are determined based on a general vulnerability rating system. When the score of a high-risk vulnerability is greater than a preset value, it is identified as a high-risk vulnerability.

6. A device for mining related terms of high-risk vulnerability components, characterized in that, The apparatus for performing the high-risk vulnerability component association word mining method as described in claim 1 includes: The dataset acquisition module is used to acquire datasets, which are information corresponding to components and vulnerabilities respectively. The knowledge graph construction module is used to construct a knowledge graph based on the dataset and a preset knowledge graph paradigm, wherein the preset knowledge graph paradigm is used to represent entities, attributes and relationships in the knowledge graph. The associated word mining module is used to mine high-risk vulnerability associated words of any component based on the dataset, the knowledge graph, and the preset multidimensional corpus, using the vulnerability severity level scoring standard and the TF-IDF algorithm. The preset multidimensional corpus consists of multiple corpora constructed according to various classification methods of the components.

7. The high-risk vulnerability component associated word mining device according to claim 6, characterized in that, The dataset includes: vulnerability articles, and the device further includes: The security information extraction module is used to extract vulnerability information from multiple security dimensions based on the vulnerability article. The indicator value determination module is used to determine the indicator values ​​of the vulnerability in multiple security dimensions based on the information of the vulnerability in multiple security dimensions. The indicator value storage module is used to store information about the vulnerability in multiple security dimensions and the indicator values ​​of the vulnerability in multiple security dimensions in a knowledge graph.

8. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the high-risk vulnerability component association word mining method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the high-risk vulnerability component association word mining method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Data processing method and device, storage medium and processor

    CN114791944A

  • Knowledge graph construction method, CWE community description method and storage medium

    CN116108847A