Construction method of multi-source heterogeneous vulnerability database
Through monitoring, collection and redundant analysis methods, the multi-source heterogeneous vulnerability database is integrated, and the problems of data dispersion and structural differences are solved, the comprehensiveness and standardization of the vulnerable database are improved, and the credibility of the data is enhanced.
Patent Information
- Application Number
- CN202510401271.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing vulnerable databases have problems such as data dispersion, structural differences and inconsistent information, which leads to limited data validity and quality, making it difficult to build a comprehensive and accurate secure database.
By monitoring commonly used vulnerability databases, official security announcements and third-party sources, the original vulnerability data is dynamically collected, and crawling technology and hash matching processing are used, and redundant analysis is carried out in combination with CVE identifiers and data source confidence, it is integrated and stored to a multi-source heterogeneous vulnerability database.
Improves the comprehensiveness, standardization and credibility of the vulnerable database, reduces data duplication, and enhances the available value of the data.
Smart Images

Figure CN120256505A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for constructing a multi-source heterogeneous vulnerability database, belonging to the technical field of software security. Background Art
[0002] Vulnerability databases are crucial in the field of software security. They collect, maintain, and disclose detailed information about discovered security vulnerabilities, including identifiers, descriptions, affected software and scope, etc., and are widely used in academia and industry. In academia, most software security research is closely related to vulnerability databases. At different stages of the vulnerability life cycle, vulnerability databases play a key role. In addition, with the rise of code reuse, vulnerability databases also support security work such as SCA, promoting research progress in the field of software security. In industry, vulnerability databases assist governments, enterprises, and third-party organizations in coping with cyber risks. The government uses them to uniformly collect and verify vulnerabilities, issue early warnings, and conduct emergency responses, improving the national security research level and prevention capabilities; enterprises use them to monitor, evaluate, and repair software security risks, reducing user losses and the risk of information leakage; the vulnerability governance and management tools of third-party agencies also rely on vulnerability databases to provide information.
[0003] Although many influential software security databases have emerged in recent years, they are managed separately by different entities, with different focuses, and there are still problems such as the fusion of multi-source heterogeneous data information, which affects the data validity and quality and restricts the research on security vulnerabilities. At the same time, there are differences in the recording of the same vulnerability information in different databases. Therefore, constructing a comprehensive and accurate security database has become an urgent problem to be solved, which also provides an important background and direction for the research and development of related patents. The current main problems are as follows: (1) Vulnerability data is relatively scattered. On the one hand, the vulnerabilities included in each vulnerability database are not completely the same. On the other hand, newly emerged vulnerabilities may not be reported to the CVE organization and are scattered in sources such as vendor security announcements that are not easily accessible to the public.
[0004] (2) There are differences in data structures. The presentation structures of the obtained vulnerability data are different. Some data sources provide downloads in formats such as JSON or xml (such as CNNVD, etc.); some data sources are displayed in hypertext markup language on the official website (such as SNYK); for some data sources, the vulnerability information is presented in a text structure mixed together.
[0005] (3) Vulnerability information is inconsistent. On the one hand, the fields of each vulnerability database are inconsistent. On the other hand, for the same field of the same vulnerability, different vulnerability databases may record inconsistent vulnerability information. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for constructing a multi-source heterogeneous vulnerability database, which can improve the data comprehensiveness, standardization, and integrity of the vulnerability database.
[0007] To achieve the above object, the present invention provides the following technical solutions: The present invention provides a method for constructing a multi-source heterogeneous vulnerability database, including: Determine the collection targets for three types of data sources, namely, common vulnerability databases, official security bulletins, and third-party sources, and monitor them to dynamically obtain original vulnerability data; Parse the structured data, semi-structured data, and unstructured data in the original vulnerability data respectively, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database. Based on the CVE identifier and the data source confidence level, perform redundancy analysis on the same vulnerability from different data sources, and store the vulnerability data after redundancy analysis in the multi-source heterogeneous vulnerability database.
[0008] Preferably, the determining the collection targets for three types of data sources, namely, common vulnerability databases, official security bulletins, and third-party sources, and monitoring them to dynamically obtain original vulnerability data includes: For the three types of data sources, namely, common vulnerability databases, official security bulletins, and third-party sources, determine the data sources that can be used as collection targets respectively according to the set selection principles; For the collection targets, use web crawler technology for automated collection and perform anti-web crawler mechanism processing to obtain original vulnerability data; Generate a hash value based on the original vulnerability data, and perform hash matching to determine whether the original vulnerability data is retained; Continuously monitor the collection targets in the above manner to dynamically obtain original vulnerability data.
[0009] Preferably, the generating a hash value based on the original vulnerability data and performing hash matching to determine whether the original vulnerability data is retained includes: Perform hash calculation on the original vulnerability data using the selected hash algorithm; Match the hash value of the newly collected original vulnerability data with the hash values of the existing vulnerability data in the multi-source heterogeneous vulnerability database one by one; if they match, it is determined that the newly collected original vulnerability data is repeated with the known vulnerabilities in the multi-source heterogeneous vulnerability database, and then discard the original vulnerability data; if it does not match all the data in the multi-source heterogeneous vulnerability database, retain the original vulnerability data.
[0010] Preferably, the continuously monitoring the collection targets to dynamically obtain original vulnerability data includes: Start the monitoring mechanism to continuously monitor the data changes of the collection targets; When the data of the collection target changes, the monitoring mechanism promptly captures the change and starts collecting the changed data of the collection target; For the collected changed data, obtain the original vulnerability data.
[0011] Preferably, perform data parsing on the structured data, semi-structured data, and unstructured data in the original vulnerability data respectively, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database, including: Design the tables and their fields of the multi-source heterogeneous vulnerability database according to the national standard documents; For the structured data in the original vulnerability data, use the JSON library or the xmltodict library to parse the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, set the value of this field to null; For the semi-structured data in the original vulnerability data, use the HTML parsing library to parse the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, set the value of this field to null; For the unstructured data in the original vulnerability data, combine the two methods of regularization matching and named entity recognition to parse the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, set the value of this field to null; The fields include: vulnerability identifier, vulnerability name, vulnerability description, vulnerability release time, author of the vulnerability report, verifier of the vulnerability report, source of the vulnerability data, original identifier of the vulnerability, related CVE identifier of the vulnerability, vulnerability type, vulnerability severity level, name of the product affected by the vulnerability, scope of vulnerability impact, related link, and existence description.
[0012] Preferably, for the unstructured data in the original vulnerability data, combining the two methods of regularization matching and named entity recognition to parse the original vulnerability data includes: Traverse each original vulnerability data, and use the named entity recognition method to extract the key information in the original vulnerability data: the name of the product affected by the vulnerability and the scope of vulnerability impact; Use the regularization matching method to extract other fields in the original vulnerability data; Extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database.
[0013] Preferably, the redundancy analysis of the same vulnerability in different data sources based on the CVE identifier and the data source confidence level, and storing the vulnerability data after redundancy analysis into a multi-source heterogeneous vulnerability database includes: Judging whether the vulnerability to be stored is the same vulnerability from different data sources as the known vulnerabilities in the multi-source heterogeneous vulnerability database based on the CVE identifier of the vulnerability; For the same vulnerability from different data sources, retain the vulnerability data with a high data source confidence level and store it in the multi-source heterogeneous vulnerability database; For vulnerabilities that are not the same vulnerability from different data sources, directly store them in the multi-source heterogeneous vulnerability database.
[0014] Preferably, judging whether the vulnerability to be stored is the same vulnerability from different data sources as the known vulnerabilities in the multi-source heterogeneous vulnerability database based on the CVE identifier of the vulnerability includes: Match the CVE identifier of the vulnerability to be stored with the CVE identifiers of the known vulnerabilities in the multi-source heterogeneous vulnerability database one by one. If the two match exactly, it is determined to be the same vulnerability from different data sources. If the two partially match or do not match at all, it is determined not to be the same vulnerability from different data sources.
[0015] Preferably, for the same vulnerability from different data sources, retaining the vulnerability data with a high data source confidence level and storing it in the multi-source heterogeneous vulnerability database includes: Measure the confidence level of the data source from three aspects: the attributes of the maintainer, the quality of the vulnerability report, and whether there is vulnerability verification data, and sort the total confidence levels of each data source; Retain the vulnerability data of the data source with the highest total confidence level, and merge the relevant links in the vulnerability data of the data sources with low confidence levels into the relevant links of the retained vulnerability data, and store it in the multi-source heterogeneous vulnerability database in a covering manner.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The method for constructing a multi-source heterogeneous vulnerability database provided by the present invention collects more scattered vulnerability data through automatic collection of multi-source data, improving the comprehensiveness and integrity of vulnerability data; through the heterogeneous data parsing method, it integrates and standardizes the vulnerability data with different structures from different data sources, and designs the database tables according to national standards, improving the standardization of vulnerability data; through the redundancy analysis method based on CVE identification and data source confidence level, it eliminates the problem of data duplication in the database, increasing the credibility and available value of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of the method for constructing a multi-source heterogeneous vulnerability database provided by an embodiment of the present invention. Detailed implementation mode
[0018] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the implementation modes and the accompanying drawings. Herein, the illustrative implementation modes of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0019] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0020] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0021] It should be emphasized here that the step marks mentioned hereinafter are not intended to limit the sequence of the steps, but should be understood that the steps can be executed in the sequence mentioned in the embodiments, or different from the sequence in the embodiments, or several steps can be executed simultaneously.
[0022] The embodiments of the present patent will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present patent and should not be construed as limiting the present patent. Without conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.
[0023] The embodiments of the present invention provide a method for constructing a multi-source heterogeneous vulnerability database. Refer to Figure 1 , which specifically includes the following steps: Step 1: Determine the collection targets for three types of data sources, namely, common vulnerability databases, official security announcements, and third-party sources, and monitor them to dynamically obtain the original vulnerability data.
[0024] Specifically, it may include the following sub-steps: Step 1.1: Define the selection principles for the three data source types of common vulnerability databases, official security bulletins, and third-party sources, and determine the data sources that can be used as collection targets according to the selection principles. Among them, common databases may include: NVD, CVE List, CNVD, CNNVD, OSV, GitHub Advisory Database, Exploit-DB, RedHat, Snyk, etc.; official security bulletins such as those released on the official websites of manufacturers such as Microsoft, Apple, and Huawei; third-party sources include security information collected and released by third-party institutions or organizations, such as opensuse, openwall, gentoo, packetstorm, etc.
[0025] Step 1.2: For the collection targets, use web crawler technology for automated collection. Since some data sources apply anti-crawler mechanisms, such as NVD, CNVD, etc., special processing needs to be carried out on the anti-crawler mechanisms in these data sources, and finally the original vulnerability data is obtained.
[0026] Step 1.3: Generate hash values based on the original vulnerability data, and perform hash matching to determine whether the vulnerability data should be retained or discarded.
[0027] It includes the following sub-steps: Step 1.3.1: Use the selected hash algorithm to calculate the hash of the collected original vulnerability data.
[0028] Step 1.3.2: Match the hash values of the newly collected original vulnerability data with the hash values of the existing vulnerability data in the multi-source heterogeneous vulnerability database one by one.
[0029] Step 1.3.3: If there is a match, it is determined that the newly collected original vulnerability data is duplicate with the known vulnerabilities in the multi-source heterogeneous vulnerability database, and the original vulnerability data is discarded and transferred to Step 1.2; if there is no match with all the data in the multi-source heterogeneous vulnerability database, the original vulnerability data is retained and transferred to Step 1.2.
[0030] Step 1.4: Monitor the collection targets and perform dynamic collection on the collection targets.
[0031] It includes the following sub-steps: Step 1.4.1: Start the monitoring mechanism and continuously monitor the data changes of the collection targets.
[0032] Step 1.4.2: When the data of the collection target changes, the monitoring mechanism will capture the changes in time and start collecting the changed data of the collection target, and jump to Step 1.2.
[0033] Step 2: Parse the structured data, semi-structured data, and unstructured data in the original vulnerability data respectively, and extract the required fields and corresponding vulnerability information.
[0034] Specifically, it may include the following sub-steps: Step 2.1: Before data parsing, it is necessary to design the tables and their fields in the multi-source heterogeneous vulnerability database to make the data format unified after parsing. According to the national standard document "Information Security Technology - Network Security Vulnerability Identification and Description Specification", the fields should at least include: vulnerability identifier, vulnerability name, vulnerability description, vulnerability release time, author of the vulnerability report, verifier of the vulnerability report, source of the vulnerability data, original identifier of the vulnerability, relevant CVE identifier of the vulnerability, vulnerability type, vulnerability severity level, name of the product affected by the vulnerability, scope of vulnerability impact, relevant link, and existence description.
[0035] Step 2.2: For the structured data in the original vulnerability data, directly parse it and extract the required fields and corresponding vulnerability information according to the table design of the multi-source heterogeneous vulnerability database. Structured data is mainly presented in JSON or xml format (such as GHSA, CNNVD, etc.), and each piece of vulnerability data consists of several fields established by the original data source. Specifically, it includes the following sub-steps: Step 2.2.1: Traverse all JSON or xml files, and use the JSON library or xmltodict library to parse the original vulnerability data.
[0036] Step 2.2.2: According to the table design of the multi-source heterogeneous vulnerability database, extract the required fields and corresponding vulnerability information. If a required field is not included in the original vulnerability data, set the value of this field to null.
[0037] Step 2.3: For the semi-structured data in the original vulnerability data, use an HTML parsing library to parse the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the table design of the multi-source heterogeneous vulnerability database. Semi-structured data is mainly presented in HTML form, and the vulnerability information exists in various tag pairs in HTML. Specifically, it includes the following sub-steps: Step 2.3.1: Traverse each piece of original vulnerability data and use parsing libraries such as beautiful soup and lxml to parse it.
[0038] Step 2.3.2: According to different data sources, analyze the vulnerability field information corresponding to each tag pair in HTML.
[0039] Step 2.3.3: Extract the required fields and the corresponding vulnerability information according to the table design of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, set the value of this field to null.
[0040] Step 2.4: For the unstructured data in the original vulnerability data, parse the original vulnerability data by combining the methods of regular expression matching and named entity recognition, and extract the required fields and the corresponding vulnerability information according to the table design of the multi-source heterogeneous vulnerability database. Specifically, it includes the following sub-steps: Step 2.4.1: Traverse each piece of original vulnerability data, and use the named entity recognition method to extract the key information in the original vulnerability data: the name of the product affected by the vulnerability and the scope of vulnerability impact.
[0041] Step 2.4.2: Use the regular expression matching method to extract other information fields in the original vulnerability data, such as: vulnerability name, vulnerability description, author of the vulnerability report, vulnerability release time, CVE identifier related to the vulnerability, vulnerability type, vulnerability severity level, and related links, etc. Step 2.4.3: Extract the required fields and the corresponding vulnerability information required by the multi-source heterogeneous vulnerability database according to the table design of the multi-source heterogeneous vulnerability database. If a required field is not included in the original data, set the value of this field.
[0042] Step 3: Perform redundancy analysis on the same vulnerability from different data sources based on the CVE identifier related to the vulnerability and the data source confidence, and store the analyzed vulnerability data in the multi-source heterogeneous vulnerability database.
[0043] Specifically, it can include the following sub-steps: Step 3.1: Judge whether the vulnerability to be stored is the same vulnerability from different data sources as the known vulnerabilities in the multi-source heterogeneous vulnerability database based on the CVE identifier related to the vulnerability. It includes the following sub-steps: Step 3.1.1: Match the CVE identifier of the vulnerability to be stored with the CVE identifiers of the known vulnerabilities in the multi-source heterogeneous vulnerability database one by one.
[0044] Step 3.1.2: If the two match exactly, it is considered that they are the same vulnerability from different data sources, and go to Step 3.2.
[0045] Step 3.1.3: If the two partially match (only some of the multiple CVE identifiers match successfully) or do not match at all, it is considered that they are not the same vulnerability from different data sources, and store them in the multi-source heterogeneous vulnerability database.
[0046] Step 3.2: Measure the confidence of data sources from three aspects: the attributes of maintainers, the quality of vulnerability reports, and the existence of vulnerability verification data. For the maintainer attribute (1 - 3 points), if the maintainer is a country, add 3 points; if the maintainer is a large enterprise or research team, add 2 points; if the maintainer is a small enterprise or other institution, add 1 point. For the vulnerability report quality attribute (1 - 3 points), let researchers with research experience in the field of security vulnerabilities review and score the vulnerability reports of each data source. For the existence of vulnerability verification data attribute (0 - 1 point), check whether there is information related to vulnerability verification in the vulnerability reports of each data source. Finally, sort the confidence of each data source.
[0047] Step 3.3: For the same vulnerability from different data sources, retain the vulnerability data with the highest data source confidence and store it in the multi-source heterogeneous vulnerability database. It includes the following sub-steps: Step 3.3.1: When the vulnerability to be stored and the known vulnerabilities in the multi-source heterogeneous vulnerability database are the same vulnerability from different data sources, compare the confidence of the data sources of the two.
[0048] Step 3.3.2: Retain the vulnerability data of the data source with the highest confidence, merge the relevant links in the vulnerability data of the data source with low confidence into the relevant links of the retained vulnerability data, and store it in the multi-source heterogeneous vulnerability database in an overwriting manner.
[0049] Step 3.3.3: Store the vulnerability data in the multi-source heterogeneous vulnerability database.
[0050] The method for constructing a multi-source heterogeneous vulnerability database provided in this embodiment collects more scattered vulnerability data through automated multi-source data collection, improving the comprehensiveness and integrity of vulnerability data; through the heterogeneous data parsing method, it integrates and standardizes the vulnerability data with different structures from different data sources, and designs the database tables according to national standards, improving the standardization of vulnerability data; through the redundancy analysis method based on CVE identification and data source confidence, it eliminates the problem of data duplication in the database, increasing the credibility and available value of the data.
[0051] Based on the same inventive concept, another embodiment of the present invention provides a device for constructing a multi-source heterogeneous vulnerability database. The device includes: A data collection module, used to monitor the collection targets of three types of data sources, namely, common vulnerability databases, official security announcements, and third-party sources, and dynamically obtain original vulnerability data; A field extraction module, used to respectively perform data parsing on the structured data, semi-structured data, and unstructured data in the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database; A data storage module, which is used to perform redundancy analysis on the same vulnerability of different data sources based on CVE identifiers and data source confidence levels, and store the vulnerability data after redundancy analysis into a multi-source heterogeneous vulnerability database.
[0052] It should be noted that the device embodiment corresponds to the above method embodiment, and the implementation manners of the above method embodiment are all applicable to the device embodiment and can achieve the same or similar technical effects, so they will not be elaborated here.
[0053] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0054] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0055] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0056] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for constructing a multi-source heterogeneous vulnerability database, characterized in that Including: Monitor the collection targets of three types of data sources, namely the common vulnerability database, official security bulletins, and third-party sources, and dynamically obtain the original vulnerability data; Parse the structured data, semi-structured data, and unstructured data in the original vulnerability data respectively, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database; Perform redundancy analysis on the same vulnerability from different data sources based on the CVE identifier and data source confidence, and store the vulnerability data after redundancy analysis in the multi-source heterogeneous vulnerability database.
2. The construction method of a multi-source heterogeneous vulnerability database according to claim 1, characterized in that, The monitoring of the collection targets of three types of data sources, namely the common vulnerability database, official security bulletins, and third-party sources, and dynamically obtaining the original vulnerability data includes: For the three types of data sources, namely the common vulnerability database, official security bulletins, and third-party sources, determine the data sources that can be used as collection targets according to the set selection principles respectively; For the collection targets, use web crawler technology for automated collection and perform anti-crawler mechanism processing to obtain the original vulnerability data; Generate a hash value based on the original vulnerability data and perform hash matching to determine whether the original vulnerability data is retained; Continuously monitor the collection targets in the above manner and dynamically obtain the original vulnerability data.
3. The construction method of a multi-source heterogeneous vulnerability database according to claim 2, characterized in that, The generating a hash value based on the original vulnerability data and performing hash matching to determine whether the original vulnerability data is retained includes: Perform hash calculation on the original vulnerability data using the selected hash algorithm; Match the hash value of the newly collected original vulnerability data with the hash values of the existing vulnerability data in the multi-source heterogeneous vulnerability database one by one; if they match, it is determined that the newly collected original vulnerability data is repeated with the known vulnerabilities in the multi-source heterogeneous vulnerability database, and then discard the original vulnerability data; if it does not match all the data in the multi-source heterogeneous vulnerability database, retain the original vulnerability data.
4. A method for constructing a multi-source heterogeneous vulnerability database according to claim 2, characterized in that, The continuous monitoring of the collection targets and dynamically obtaining the original vulnerability data includes: Start the monitoring mechanism and continuously monitor the data changes of the collection targets; When the data of the collection target changes, the monitoring mechanism promptly captures the change and starts to collect the changed data of the collection target; For the collected changed data, obtain the original vulnerability data.
5. A method for constructing a multi-source heterogeneous vulnerability database according to claim 1, characterized in that The parsing of the structured data, semi-structured data, and unstructured data in the original vulnerability data respectively, and extracting the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database includes: Design the tables and their fields of the multi-source heterogeneous vulnerability database according to the national standard documents; For the structured data in the original vulnerability data, use the JSON library or the xmltodict library to parse the original vulnerability data, and extract the required fields and corresponding vulnerability information according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, set the value of this field to null; For the semi-structured data in the original vulnerability data, the original vulnerability data is parsed with the help of an HTML parsing library, and the required fields and corresponding vulnerability information are extracted according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, the value of this field is set to null; For the unstructured data in the original vulnerability data, the original vulnerability data is parsed by combining two methods: regularization matching and named entity recognition, and the required fields and corresponding vulnerability information are extracted according to the designed tables of the multi-source heterogeneous vulnerability database. If a required field is not included in the original vulnerability data, the value of this field is set to null; The fields include: vulnerability identifier, vulnerability name, vulnerability description, vulnerability release time, author of the vulnerability report, verifier of the vulnerability report, source of the vulnerability data, original identifier of the vulnerability, related CVE identifier of the vulnerability, vulnerability type, vulnerability severity level, name of the product affected by the vulnerability, scope of vulnerability impact, related link, and existence description.
6. The construction method of a multi-source heterogeneous vulnerability database according to claim 5, characterized in that, The parsing of the unstructured data in the original vulnerability data by combining two methods: regularization matching and named entity recognition includes: Traverse each piece of original vulnerability data, and use the named entity recognition method to extract the key information in the original vulnerability data: the name of the product affected by the vulnerability and the scope of vulnerability impact; Use the regularization matching method to extract other fields from the original vulnerability data; According to the designed tables of the multi-source heterogeneous vulnerability database, extract the required fields and corresponding vulnerability information.
7. A method for constructing a multi-source heterogeneous vulnerability database according to claim 1, characterized in that The redundancy analysis of the same vulnerability from different data sources based on the CVE identifier and data source confidence level, and storing the vulnerability data after redundancy analysis into the multi-source heterogeneous vulnerability database includes: Based on the CVE identifier of the vulnerability, judge whether the vulnerability to be stored is the same vulnerability from different data sources as the known vulnerabilities in the multi-source heterogeneous vulnerability database; For the same vulnerability from different data sources, retain the vulnerability data with a high data source confidence level and store it in the multi-source heterogeneous vulnerability database; For vulnerabilities that are not the same vulnerability from different data sources, directly store them in the multi-source heterogeneous vulnerability database.
8. The construction method of a multi-source heterogeneous vulnerability database according to claim 7, characterized in that The judgment of whether the vulnerability to be stored is the same vulnerability from different data sources as the known vulnerabilities in the multi-source heterogeneous vulnerability database based on the CVE identifier of the vulnerability includes: Match the CVE identifier of the vulnerability to be stored with the CVE identifiers of the known vulnerabilities in the multi-source heterogeneous vulnerability database one by one. If the two match exactly, it is determined to be the same vulnerability from different data sources. If the two partially match or do not match at all, it is determined not to be the same vulnerability from different data sources.
9. The construction method of a multi-source heterogeneous vulnerability database according to claim 7, characterized in that The retention of the vulnerability data with a high data source confidence level and storing it in the multi-source heterogeneous vulnerability database for the same vulnerability from different data sources includes: Measure the confidence level of the data source from three aspects: the attributes of the maintainer, the quality of the vulnerability report, and whether there is vulnerability verification data, and sort the total confidence levels of each data source; Retain the vulnerability data of the data source with the highest overall confidence, and merge the relevant links in the vulnerability data of the data sources with low confidence into the relevant links of the retained vulnerability data, and store them in a multi-source heterogeneous vulnerability database in an overwriting manner.
Citation Information
Cited By
Multi-source vulnerability data fusion method and related product
CN121071883A