Fusion processing method and device for multi-source vulnerability data, equipment and medium

By dividing multi-level knowledge base identification and mapping relationships to process multi-source vulnerability data, the false alarm and missed response problems caused by different scanning results of different channels are solved, and more accurate vulnerability data integration and network security improvement are achieved.

CN120277624AActive Publication Date: 2025-07-08PENG CHENG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510766045.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Due to the different vulnerability scanning performance of different pathways, there are significant differences in the scanning results of multi-source pathways under the same network environment, which is difficult to accurately integrate, increasing the probability of false alarms or missed reports of vulnerable data and reducing network security.

Method used

By dividing the hierarchical order of multi-level knowledge base identifiers, matching the knowledge base identifier of vulnerability data in sequence, and determining whether there is a mapping relationship in the mapping knowledge base, integrating low-level vulnerability data with high-level vulnerability data, using a similarity algorithm to process non-standard vulnerability data, establishing vulnerability fusion tables and mapping knowledge bases, and achieving effective integration of multi-source vulnerability data.

Benefits of technology

It reduces the probability of false alarms or misreports of vulnerable data caused by significant differences in scanning results of different channels, and improves network security and management accuracy of vulnerable data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277624A_ABST
    Figure CN120277624A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a fusion processing method and device for multi-source vulnerability data, equipment and a medium, multiple pieces of vulnerability data found through multiple different ways are obtained, fusion processing is carried out on the vulnerability data in sequence, and for the first vulnerability data currently subjected to fusion processing, if the first vulnerability data carries a corresponding first knowledge base identifier, the first knowledge base identifier corresponding to the first vulnerability data is identified. Matching the first knowledge base identifiers in sequence from high to low according to the level sequence of the pre-arranged multi-level knowledge base identifiers; when the first knowledge base identifier is not matched with any first type of identifiers but is matched with one of the second type of identifiers, determining whether there is a first type of identifier having a mapping relationship with the first knowledge base identifier in a mapping knowledge base, and if there is a first type of identifier having a mapping relationship with the first knowledge base identifier, taking the first type of identifier having the mapping relationship as a target knowledge base identifier; and distributing the first vulnerability data to the target knowledge base identifier to obtain a corresponding vulnerability fusion result, thereby effectively integrating multi-source data, and reducing the probability of false report or missing report of the vulnerability data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security technologies, and in particular, to a method, apparatus, device, and medium for fusing and processing multi-source vulnerability data. Background Art

[0002] With the wide popularization of the Internet and the rapid development of computer technologies, information technology has penetrated into all fields of society, greatly enriching and simplifying people's lives. However, network security issues have become increasingly prominent, and network attack means have been constantly renovated, seriously threatening the information security of individuals and organizations. Therefore, how to accurately identify vulnerability data is crucial for improving network security. Currently, vulnerability data is often scanned through various channels.

[0003] In related technologies, due to the limitations of a single channel in vulnerability scanning, the implementation of multi-source vulnerability data scanning based on multiple different channels has gradually become a development trend. However, due to the different vulnerability scanning performances of different channels and the different vulnerability knowledge bases or scanning means they rely on, there are significant differences in the scanning results of multi-source channels in the same network environment and it is difficult to accurately fuse them, resulting in the inability to effectively integrate the vulnerability data collected from multi-source channels, thereby increasing the probability of false positives or false negatives of vulnerability data and ultimately reducing network security. Summary of the Invention

[0004] The main objective of the embodiments of the present disclosure is to propose a method, apparatus, device, and medium for fusing and processing multi-source vulnerability data, which can effectively integrate data from multi-source channels, reduce the probability of false positives or false negatives of vulnerability data, and ultimately improve network security.

[0005] To achieve the above objective, a first aspect of the embodiments of the present disclosure proposes a method for fusing and processing multi-source vulnerability data, including: Obtaining multiple vulnerability data discovered by multiple different channels respectively, where when the vulnerability data carries a knowledge base identifier of a corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; Sequentially performing fusion processing on each of the vulnerability data, and for the first vulnerability data currently undergoing fusion processing, if the first vulnerability data carries a corresponding first knowledge base identifier, matching the first knowledge base identifier in descending order according to the hierarchical order of the pre-arranged multi-level knowledge base identifiers, where the multi-level knowledge base identifiers include multiple first-type identifiers and multiple second-type identifiers, and the level of the first-type identifiers is higher than that of the second-type identifiers; When the first knowledge base identifier does not match any of the first type of identifiers but matches one of the second type of identifiers, determine whether there is a first type of identifier having a mapping relationship with the first knowledge base identifier in a preset mapping knowledge base; When there is a first type of identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the first type of identifier having the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

[0006] In some embodiments, after successively performing fusion processing on each of the vulnerability data and for the first vulnerability data currently undergoing fusion processing, the multi-source vulnerability data fusion processing method further includes: If the first vulnerability data does not carry a corresponding first knowledge base identifier, obtain a plurality of second vulnerability data from a preset vulnerability fusion table, where the second knowledge base identifier of the second vulnerability data does not belong to the first type of identifier nor the second type of identifier; Determine the vulnerability similarity between the first vulnerability data and each of the second vulnerability data to obtain corresponding similarity values; When any one of the similarity values is greater than or equal to a preset first similarity threshold, determine the second knowledge base identifier of the similar second vulnerability data as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

[0007] In some embodiments, the determining the vulnerability similarity between the first vulnerability data and each of the second vulnerability data to obtain corresponding similarity values includes: Respectively determine the vulnerability attributes of a plurality of different attribute types under the first vulnerability data and the second vulnerability data, and configure corresponding similarity algorithms for different attribute types; For each of the first vulnerability data and each of the second vulnerability data, under the corresponding similarity algorithm, successively determine the attribute similarity between the vulnerability attributes under the same attribute type to obtain sub-similarity values under a plurality of attribute types; Based on the weights under each attribute type, perform weighted calculation on a plurality of the sub-similarity values to obtain the similarity value between the first vulnerability data and each of the second vulnerability data.

[0008] In some embodiments, the successively determining the attribute similarity between the vulnerability attributes under the same attribute type to obtain sub-similarity values under a plurality of attribute types includes: Calculate the edit distance and the first Jaccard similarity value between the vulnerability attributes whose attribute type is component name, and determine the sub-similarity value under the component name by combining the edit distance and the first Jaccard similarity value; Calculate the hierarchical version split vector and the first cosine similarity value between the vulnerability attributes whose attribute type is component version, and determine the sub-similarity value under the component version by combining the hierarchical version split vector and the first cosine similarity value; Calculate the first domain-enhanced semantic vector feature and the second cosine similarity value between the vulnerability attributes whose attribute type is vulnerability description, and determine the sub-similarity value under the vulnerability description by combining the first domain-enhanced semantic vector feature and the second cosine similarity value; Calculate the action and object semantic role annotation result and the third cosine similarity value between the vulnerability attributes whose attribute type is repair suggestion, and determine the sub-similarity value under the repair suggestion by combining the annotation result and the third cosine similarity value; Calculate the classification tree path distance or the second domain-enhanced semantic vector feature between the vulnerability attributes whose attribute type is vulnerability type, and determine the sub-similarity value under the vulnerability type by combining the classification tree path distance or the second domain-enhanced semantic vector feature; Calculate the TF-IDF vector feature and the fourth cosine similarity value between the vulnerability attributes whose attribute type is exploitation method, and determine the sub-similarity value under the exploitation method by combining the TF-IDF vector feature and the fourth cosine similarity value; Calculate the second Jaccard similarity value of the asset type between the vulnerability attributes whose attribute type is impact scope, and determine the sub-similarity value under the impact scope based on the second Jaccard similarity value.

[0009] In some embodiments, after determining whether there is a first type of identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge base identifier, the multi-source vulnerability data fusion processing method further includes: When there is no first type of identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, use the second type of identifier that matches the first knowledge base identifier as the target knowledge base identifier, and assign the first vulnerability data to under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0010] In some embodiments, the multiple second type of identifiers include multiple second type of identifiers under the first vulnerability knowledge base, and multiple second type of identifiers under the second vulnerability knowledge base at the same level as the first vulnerability knowledge base; After the first type of identifier having a mapping relationship with the first knowledge base identifier does not exist in the mapping knowledge base, the multi-source vulnerability data fusion processing method further includes: When the first knowledge base identifier matches the second type of identifier under one of the first vulnerability knowledge bases, determine whether there is a second type of identifier under the second vulnerability knowledge base having a mapping relationship with the first knowledge base identifier in the mapping knowledge base; When there is a second type of identifier under the second vulnerability knowledge base having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the second type of identifier under the second vulnerability knowledge base having the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

[0011] In some embodiments, after sequentially matching the first knowledge base identifier from high to low according to the level order of the pre-arranged multi-level knowledge base identifiers, the multi-source vulnerability data fusion processing method further includes: When the first knowledge base identifier matches one of the first type of identifiers, use the first type of identifier that matches the first knowledge base identifier as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

[0012] In some embodiments, the allocating the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result includes: Determine whether there is a vulnerability storage record of the target knowledge base identifier in the preset vulnerability fusion table; When there is a vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table, update the vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table based on the first vulnerability data to obtain a corresponding vulnerability fusion result; When there is no vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table, obtain the vulnerability description information of the first vulnerability data from the vulnerability knowledge base corresponding to the target knowledge base identifier, and create a vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table based on the vulnerability description information to obtain a corresponding vulnerability fusion result.

[0013] In some embodiments, the multi-source vulnerability data fusion processing method further includes: Determine the vulnerability similarity between the vulnerability data under any one of the first type of identifiers and the vulnerability data under any one of the second type of identifiers to obtain a corresponding similarity value; When any one of the similarity values is greater than or equal to a preset second similarity threshold, a mapping relationship between the corresponding first type of identifier and the second type of identifier is established for the determined similar vulnerability data and saved in the mapping knowledge base.

[0014] To achieve the above object, a second aspect of the embodiments of the present disclosure provides a multi-source vulnerability data fusion processing device, including: A multi-source data acquisition module, configured to respectively acquire a plurality of vulnerability data discovered by a plurality of different channels. Among them, when the vulnerability data carries a knowledge base identifier of a corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the plurality of vulnerability knowledge bases; An identifier matching module, configured to sequentially perform fusion processing on each of the vulnerability data, and for the first vulnerability data currently undergoing fusion processing, if the first vulnerability data carries a corresponding first knowledge base identifier, sequentially match the first knowledge base identifier from high to low according to the hierarchical order of the pre-arranged multi-level knowledge base identifiers, where the multi-level knowledge base identifiers include a plurality of first type of identifiers and a plurality of second type of identifiers, and the level of the first type of identifier is higher than that of the second type of identifier; A mapping determination module, configured to determine whether there is a first type of identifier having a mapping relationship with the first knowledge base identifier in a preset mapping knowledge base when the first knowledge base identifier does not match any one of the first type of identifiers but matches one of the second type of identifiers; A result determination module, configured to use the first type of identifier having a mapping relationship as a target knowledge base identifier and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result when there is a first type of identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base.

[0015] To achieve the above object, a third aspect of the embodiments of the present disclosure provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the multi-source vulnerability data fusion processing method described in the first aspect of the above embodiments.

[0016] To achieve the above object, a fourth aspect of the embodiments of the present disclosure provides a storage medium, where the storage medium is a computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-source vulnerability data fusion processing method described in the first aspect of the above embodiments.

[0017] In the embodiments of the present disclosure, by executing a method for fusing multi-source vulnerability data, multiple vulnerability data discovered through multiple different channels can be obtained respectively. Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; the fusion process is performed on each vulnerability data in turn, and for the first vulnerability data currently undergoing the fusion process, if the first vulnerability data carries the corresponding first knowledge base identifier, in accordance with the hierarchical order of the pre-arranged multi-level knowledge base identifiers, the first knowledge base identifier of the first vulnerability data is matched from high to low in turn. Among them, the multi-level knowledge base identifiers include multiple first-type identifiers and multiple second-type identifiers, and the level of the first-type identifier is higher than that of the second-type identifier; when the first knowledge base identifier does not match any of the first-type identifiers but matches one of the second-type identifiers, it is determined whether there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the preset mapping knowledge base; when there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, the first-type identifier having the mapping relationship is used as the target knowledge base identifier, and the first vulnerability data is assigned under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0018] Thus, in the embodiments of the present disclosure, after receiving multiple vulnerability data discovered through different channels, the fusion process can be performed on each vulnerability data respectively. During the process of processing the current first vulnerability data, if the first vulnerability data carries the corresponding first knowledge base identifier, since the embodiments of the present disclosure have pre-divided the hierarchical order of the multi-level knowledge base identifiers and defined that the level of the first-type identifier is higher than that of the second-type identifier, that is, the vulnerability knowledge base level corresponding to the first-type identifier is higher than the vulnerability knowledge base level corresponding to the second-type identifier. During the process of matching the first knowledge base identifier from high to low in accordance with the hierarchical order of the multi-level knowledge base identifiers, it can be first determined whether the first vulnerability data belongs to the first-type identifier, and then it is determined whether it belongs to the second-type identifier after it does not belong to the first-type identifier. Then, after the first vulnerability data belongs to the second-type identifier, it is determined whether there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the preset mapping knowledge base, so as to determine whether the first vulnerability data is the same as the vulnerability data indicated by the first-type identifier. Once there is a first-type identifier having a mapping relationship, it means that although the first vulnerability data is scanned through the lower-level vulnerability knowledge base corresponding to the second-type identifier, it is the same as the vulnerability data marked by the higher-level vulnerability knowledge base corresponding to the first-type identifier. Then, during the fusion process, the first-type identifier having the mapping relationship can be used as the target knowledge base identifier to achieve the assignment of the first vulnerability data under the higher-level target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0019] Compared with the solutions in the related art, the embodiments of the present disclosure can classify the levels of different vulnerability knowledge bases, and when there are differences in the vulnerability data scanned through different channels, identify whether the vulnerability data scanned at a lower level is the same as the vulnerability data under a higher-level identifier, so as to effectively integrate multi-source vulnerability data, reduce the probability of false alarms or missed reports of vulnerability data caused by significant differences in scanning results through different channels, and ultimately improve network security. Description of the Drawings

[0020] Figure 1 It is a schematic diagram of an application environment of a method for fusing and processing multi-source vulnerability data provided by an embodiment of the present disclosure; Figure 2 It is a schematic flowchart of a method for fusing and processing multi-source vulnerability data provided by an embodiment of the present disclosure; Figure 3 It is a schematic diagram of a scenario for fusing and processing multi-source vulnerability data provided by an embodiment of the present disclosure; Figure 4 It is Figure 2 A schematic flowchart further included after step 202 in Figure 5 It is Figure 4 A schematic flowchart further included in step 302 in Figure 6 It is Figure 5 A schematic flowchart further included in step 402 in Figure 7 It is a schematic flowchart further included after step 601 in an embodiment of the present disclosure; Figure 8 It is Figure 2 Another schematic flowchart further included after step 202 in Figure 9 It is Figure 2 A schematic flowchart further included in step 204 in Figure 10 It is another schematic flowchart of a method for fusing and processing multi-source vulnerability data provided by an embodiment of the present disclosure; Figure 11 It is a schematic diagram of functional modules of a device for fusing and processing multi-source vulnerability data provided by an embodiment of the present disclosure; Figure 12 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments

[0021] To enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.

[0022] It can be understood that in the specific implementation of the present disclosure, it involves retrieving initial timing data, initial sample timing data, and related data. When the above embodiments of the present disclosure are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards.

[0023] In addition, when the embodiments of the present disclosure need to retrieve initial timing data, initial sample timing data, and related data, a separate permission or separate consent for the initial timing data, initial sample timing data, and related data will be obtained by means of a pop-up window or jumping to a confirmation page, etc. After clearly obtaining the separate permission or separate consent for the initial timing data, initial sample timing data, and related data, the necessary initial timing data, initial sample timing data, and related data for the normal operation of the embodiments of the present disclosure will be retrieved.

[0024] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0025] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations: Common Vulnerabilities & Exposures (CVE) is a global vulnerability dictionary project, an open list or database, aiming to assign unique identifiers to publicly disclosed cybersecurity vulnerabilities. It provides standardized names for various publicly known information security vulnerabilities and risks. The identifier marked for vulnerability data in CVE is denoted as CVE ID.

[0026] The China National Vulnerability Database of Information Security (CNNVD) is a national information security vulnerability data management platform built and operated in China to effectively perform the functions of vulnerability analysis and risk assessment, aiming to provide services for China's information security assurance. The identifier marked for vulnerability data in CNNVD is denoted as CNNVD ID.

[0027] The China National Vulnerability Database (CNVD) is an information security vulnerability information sharing knowledge base established in China. Its main goal is to improve the overall research level and timely prevention ability in security vulnerabilities in China, and drive the development of domestic related security products. The identifier marked for vulnerability data in CNVD is denoted as CNVD ID.

[0028] With the wide popularity of the Internet and the rapid development of computer technology, information technology has penetrated into all fields of society, greatly enriching and simplifying people's lives. However, network security issues have become increasingly prominent, and network attack methods have been constantly renovated, seriously threatening the information security of individuals and organizations. Therefore, how to accurately identify vulnerability data is crucial for improving network security, and currently, vulnerability data is often scanned through various channels.

[0029] In related technologies, due to the limitations of a single channel in vulnerability scanning, the scanning of multi-source vulnerability data based on multiple different channels has gradually become a development trend. However, due to the different vulnerability scanning performances of different channels and the different vulnerability knowledge bases or scanning means they rely on, there are significant differences in the scanning results of multi-source channels in the same network environment. For example, different vulnerability knowledge bases relied on by different vulnerability scanning tools lead to differences in the scanned vulnerability data, and it is difficult to accurately integrate them, resulting in the inability to effectively integrate the vulnerability data collected from multi-source channels, thereby increasing the probability of false positives or false negatives of vulnerability data and ultimately reducing network security.

[0030] Embodiments of the present disclosure propose a method, device, equipment, and medium for fusing and processing multi-source vulnerability data to effectively integrate data from multi-source channels, reduce the probability of false positives or false negatives of vulnerability data, and ultimately improve network security.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the scenario of the implementation environment of the method for fusing and processing multi-source vulnerability data provided by embodiments of the present disclosure, including: terminal 101 and server 102.

[0032] Exemplarily, the server 102 may obtain multiple vulnerability data discovered by multiple different channels from the terminal 101 respectively. Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; the fusion process is performed on each vulnerability data in turn, and for the first vulnerability data currently undergoing the fusion process, if the first vulnerability data carries the corresponding first knowledge base identifier, in accordance with the hierarchical order of the pre-arranged multi-level knowledge base identifiers, the first knowledge base identifier of the first vulnerability data is matched from high to low. Among them, the multi-level knowledge base identifiers include multiple first-type identifiers and multiple second-type identifiers, and the level of the first-type identifiers is higher than that of the second-type identifiers; when the first knowledge base identifier does not match any of the first-type identifiers but matches one of the second-type identifiers, it is determined whether there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the preset mapping knowledge base; when there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, the first-type identifier having the mapping relationship is used as the target knowledge base identifier, and the first vulnerability data is assigned under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0033] The terminal 101 may be a mobile phone, a computer, an intelligent voice interaction device, an intelligent wearable device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 may also independently execute the multi-source vulnerability data fusion processing method. The terminal 101 and the server 102 may be directly or indirectly connected through a wired or wireless communication method, and the embodiments of the present disclosure do not limit this here.

[0034] The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, the server 102 may also be a node server in a blockchain network.

[0035] It should be noted that Figure 1 The schematic diagram of the scenario of the multi-source vulnerability data fusion processing method implementation environment shown is only an example. The scenarios described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those of ordinary skill in the art know that with the evolution of technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0036] Please refer toFigure 2 , Figure 2 is a schematic flowchart of a method for processing and fusing multi-source vulnerability data provided by an embodiment of the present disclosure. The method for processing and fusing multi-source vulnerability data can be applied to the server in the above embodiment, or jointly executed by a terminal and a server. The method for processing and fusing multi-source vulnerability data includes steps 201 to 204: Step 201: Obtain multiple vulnerability data discovered through multiple different channels respectively; Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; Step 202: Perform fusion processing on each vulnerability data in sequence. For the first vulnerability data currently undergoing fusion processing, if the first vulnerability data carries the corresponding first knowledge base identifier, match the first knowledge base identifier in descending order according to the hierarchical order of the pre-arranged multi-level knowledge base identifiers; Among them, the multi-level knowledge base identifiers include multiple first-type identifiers and multiple second-type identifiers, and the level of the first-type identifiers is higher than that of the second-type identifiers; Step 203: When the first knowledge base identifier does not match any of the first-type identifiers but matches one of the second-type identifiers, determine whether there is a first-type identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge base identifier; Step 204: When there is a first-type identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, use the first-type identifier with the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0037] Regarding the above step 201, there can be multiple channels. For example, it can be penetration testing, static and dynamic analysis, network traffic analysis, vulnerability information exchange platforms, vulnerability scanning tools, etc. In the subsequent embodiments of the present disclosure, the channel is taken as an example of a vulnerability scanning tool for illustration, but it does not represent a limitation on the channel.

[0038] A vulnerability scanning tool is a software tool used to detect security vulnerabilities existing in computer systems, network devices, or applications. Different vulnerability scanning tools may adopt different scanning techniques and algorithms to discover various types of vulnerabilities. A vulnerability knowledge base is a database storing information on various known security vulnerabilities, including detailed descriptions of vulnerabilities, scope of impact, repair suggestions, etc. Different vulnerability scanning tools may rely on different vulnerability knowledge bases for vulnerability detection. Vulnerability data is the relevant information about security vulnerabilities existing in a system or application discovered by a vulnerability scanning tool during the scanning process. A knowledge base identifier is a symbol or number used to uniquely identify each vulnerability knowledge base. Through the knowledge base identifier, different vulnerability knowledge bases on which the vulnerability data depends can be distinguished, facilitating subsequent classification and processing of the vulnerability data.

[0039] It should be noted that in the field of network security, a single vulnerability scanning tool has limitations because different tools may focus on different types of vulnerabilities or use different detection methods. By using multiple different vulnerability scanning tools, security vulnerabilities in a system can be discovered more comprehensively. Each vulnerability scanning tool scans based on a different vulnerability knowledge base, and the obtained vulnerability data carries the corresponding knowledge base identifier. The advantage of this is that when processing the vulnerability data subsequently, it can be clearly known which knowledge base each piece of vulnerability data is based on, thus providing a basis for subsequent fusion processing.

[0040] For example, assume that three different vulnerability scanning tools, such as Tool A, Tool B, and Tool C, are used to scan a terminal system for vulnerabilities. Tool A scans based on the vulnerability knowledge base K1 and discovers vulnerability a in the system, and this vulnerability data will carry the knowledge base identifier K1; Tool B scans based on the vulnerability knowledge base K2 and detects vulnerability b in the system, and this vulnerability data carries the knowledge base identifier K2; Tool C scans based on the vulnerability knowledge base K3 and discovers vulnerability c in the system, and its vulnerability data carries the knowledge base identifier K3.

[0041] It should be noted that each vulnerability scanning tool can, based on actual needs, deploy different vulnerability knowledge bases for vulnerability scanning, and moreover, the same vulnerability scanning tool can also scan based on multiple different vulnerability knowledge bases. The embodiments of the present disclosure are not limited to the scanning forms and contents of vulnerability scanning tools. When the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, it indicates that the vulnerability data is a standard vulnerability scanned by a certain path based on the corresponding vulnerability knowledge base, such as scanned by a certain vulnerability scanning tool based on the corresponding vulnerability knowledge base.

[0042] Regarding the above step 202, the first vulnerability data is the specific vulnerability data being currently processed when multiple vulnerability data are fused. It is the current operation object in the entire fusion process. The first knowledge base identifier is the knowledge base identifier of the corresponding vulnerability knowledge base carried by the first vulnerability data, which is used to indicate from which vulnerability knowledge base the vulnerability data is scanned. It should be noted that in the embodiments of the present disclosure, the fusion process can be performed on each vulnerability data in sequence. Each processed vulnerability data can be used as the first vulnerability data. After identifying that the first vulnerability data carries the corresponding first knowledge base identifier, the first knowledge base identifier of the first vulnerability data is matched in descending order according to the pre-arranged hierarchical order of the multi-level knowledge base identifiers.

[0043] The multi-level knowledge base identifier is for more effectively processing the results scanned by different vulnerability scanning tools based on different vulnerability knowledge bases. All relevant knowledge base identifiers are pre-arranged hierarchically to form a multi-level knowledge base identifier system. The first type of identifier and the second type of identifier are two different hierarchical categories in the multi-level knowledge base identifier. There are multiple first type of identifiers and multiple second type of identifiers. The vulnerability knowledge base corresponding to the first type of identifier has a higher level and has a more preferential status in the fusion process of vulnerability data; the vulnerability knowledge base corresponding to the second type of identifier has a relatively lower level.

[0044] It should be noted that different vulnerability scanning tools rely on different vulnerability knowledge bases, and there may be differences in the authority, accuracy, and coverage of these knowledge bases. By dividing the hierarchical order of the multi-level knowledge base identifiers and matching the knowledge base identifiers of the vulnerability data in descending order, the vulnerability data can be preferentially associated with the high-level knowledge bases. The advantage of this is that it can classify and integrate the vulnerability data more accurately. Because high-level knowledge bases often have higher authority and accuracy, if a vulnerability data can be matched with the first type of identifier at a high level, then it can be classified into the category corresponding to the high-level knowledge base, improving the accuracy and reliability of the vulnerability data classification.

[0045] For example, taking the case where the vulnerability data is scanned based on CVE, CNNVD, and CNVD as an example, the CVE ID, CNNVD ID, and CNVD ID form a multi-level knowledge base identifier. Among them, since CVE is more widely used and more authoritative internationally, multiple identifiers under CVE can be defined as the first type of identifier, that is, all CVE IDs are the first type of identifier, while multiple identifiers under CNNVD and CNVD are defined as the second type of identifier, that is, all CNNVD IDs and CNVD IDs are the second type of identifier.

[0046] It should be noted that, in addition to the above-mentioned examples of CVE, CNNVD, and CNVD, when the vulnerability data in the embodiments of the present disclosure is scanned based on other vulnerability knowledge bases, the embodiments of the present disclosure can select one or more knowledge bases with higher generality, uniqueness, efficiency, and compatibility from multiple vulnerability knowledge bases as high-level knowledge bases, and use the identifiers under the high-level knowledge bases as the first type of identifiers. On the contrary, the remaining knowledge bases are determined as low-level knowledge bases, and the identifiers under the low-level knowledge bases are used as the second type of identifiers to ensure comprehensive, accurate, and efficient vulnerability fusion. The embodiments of the present disclosure do not make specific limitations on this.

[0047] Regarding the above step 203, the mapping knowledge base is a pre-set knowledge base that records the mapping relationships between different vulnerability knowledge base identifiers. Specifically, in the mapping knowledge base of the embodiments of the present disclosure, the mapping relationship between the second type of identifiers and the first type of identifiers is stored, which is used to determine whether the vulnerability data scanned through the second type of identifiers is related to the high-level vulnerability data represented by the first type of identifiers. When there is a mapping relationship between a certain second type of identifier and a certain first type of identifier stored in the mapping knowledge base, it indicates that the corresponding vulnerabilities are actually the same vulnerability data. It's just that due to the different vulnerability knowledge bases they rely on, the identifiers are different, resulting in significant differences in the scanning results of different vulnerability scanning tools even for the same vulnerability data. Further, the mapping knowledge base can also store the corresponding relationships between other knowledge base identifiers, that is, the mapping knowledge base is used to store the mapping relationships of the same vulnerability data but with multiple different knowledge base identifiers.

[0048] When the embodiments of the present disclosure match the first knowledge base identifier of the first vulnerability data and find that it does not match any of the first type of identifiers but matches a certain second type of identifier, that is to say, the currently processed first vulnerability data is scanned by a vulnerability knowledge base corresponding to a low-level second type of identifier. After meeting the above preconditions, the embodiments of the present disclosure will search the pre-set mapping knowledge base to find out whether there is a first type of identifier that has a mapping relationship with the first knowledge base identifier. By searching the mapping knowledge base, it can be determined whether the vulnerability data scanned through the low-level second type of knowledge base is essentially the same vulnerability as the one scanned through the high-level first type of knowledge base.

[0049] It should be noted that different vulnerability scanning tools rely on different vulnerability knowledge bases, and the coverage and focus of these knowledge bases vary. Although the vulnerability knowledge base corresponding to the first type of identifier has a higher level and is more authoritative, some vulnerability scanning tools may not necessarily perform vulnerability scanning based on a high-level vulnerability knowledge base. Therefore, vulnerability data may be scanned in the lower-level vulnerability knowledge base corresponding to the second type of identifier, resulting in the knowledge base identifier carried by the vulnerability data belonging to the second type of identifier. Through the preset mapping knowledge base, it is possible to find out whether there is a potential association between the vulnerability data scanned in the lower-level vulnerability knowledge base and the vulnerability data marked in the high-level vulnerability knowledge base. This helps to accurately classify the vulnerability data that originally belongs to the high-level knowledge base category but is scanned by the lower-level knowledge base under the high-level knowledge base when fusing multi-source vulnerability data, thereby improving the integration effect of vulnerability data, reducing the probability of false positives and false negatives, and enhancing the accuracy of network security assessment.

[0050] Regarding step 204 above, in the preset mapping knowledge base, the first type of identifier that has a mapping relationship with the first knowledge base identifier of the current first vulnerability data is called the target knowledge base identifier. It is the final attribution identifier determined for this vulnerability data when processing vulnerability data fusion, and is used to accurately allocate the vulnerability data to the corresponding high-level vulnerability knowledge base category.

[0051] When a first type of identifier that has a mapping relationship with the first knowledge base identifier is found in the preset mapping knowledge base, in the embodiments of the present disclosure, this first type of identifier is determined as the target knowledge base identifier. Then, the first vulnerability data being processed is allocated to the category corresponding to this target knowledge base identifier, thereby obtaining a vulnerability fusion result after fusion processing.

[0052] It should be noted that due to the different vulnerability knowledge bases relied on by different vulnerability scanning tools, their scanning results may vary. Through the previous steps, it has been determined that some vulnerability data scanned through the lower-level vulnerability knowledge base is actually the same as the vulnerability data marked in the high-level vulnerability knowledge base. In this step, the vulnerability data scanned by these lower-level knowledge bases is reallocated to the high-level target knowledge base identifier, so that all vulnerability data can be classified and managed in a more accurate and reasonable manner, thereby improving the quality and usability of vulnerability data and providing a more reliable basis for subsequent network security analysis and decision-making.

[0053] Further, in the embodiments of the present disclosure, a vulnerability fusion table may be established and maintained. The vulnerability fusion table may store the determined target knowledge base identifier for each vulnerability data, and the finally formed vulnerability fusion result. Specifically, in the embodiments of the present disclosure, the vulnerability scanning tool corresponding to the first vulnerability data may also be recorded in the vulnerability fusion table. That is, under each knowledge base identifier in the vulnerability fusion table, the corresponding vulnerability scanning tool may be marked, so that it can be known through looking up the table which vulnerability scanning tools scanned the vulnerability data under different knowledge base identifiers, thus facilitating subsequent network security analysis. The embodiments of the present disclosure do not make specific limitations on this.

[0054] Next, in combination with specific application scenarios, the embodiments of the present disclosure will be illustrated by examples: Please refer to Figure 3 , Figure 3 which is a schematic diagram of the fusion processing scenario of multi-source vulnerability data provided by the embodiments of the present disclosure. In this embodiment, three vulnerability scanning tools are set, namely tool A, tool B, and tool C. Tool A depends on CVE for vulnerability scanning and obtains vulnerability a (vulnerability data), tool B depends on CNNVD for vulnerability scanning and obtains vulnerability b (vulnerability data), and tool C depends on CNVD for vulnerability scanning and obtains vulnerability c (vulnerability data). Each vulnerability data carries a corresponding knowledge base identifier. After all three tools send the collected vulnerability data to the multi-source vulnerability data fusion processing system, the multi-source vulnerability data fusion processing system can implement the fusion processing process of multi-source vulnerability data by executing the multi-source vulnerability data fusion processing method in the above embodiments.

[0055] Specifically, after the multi-source vulnerability data fusion processing system matches each vulnerability data in the order of CVE ID to CNNVD ID or CNVD ID, it is found that these three vulnerability data are indeed scanned by different vulnerability scanning tools under different vulnerability knowledge bases. However, through discovery in the mapping knowledge base, it is found that vulnerability b and vulnerability a are the same vulnerability. Therefore, finally, under the CVE ID corresponding to vulnerability a, the vulnerability a and vulnerability b scanned by tool A and tool B respectively may be recorded, that is, it is marked that the vulnerability under this CVE ID is scanned by tool A and tool B, and for the CNVD ID corresponding to vulnerability c, the vulnerability c scanned by tool C is recorded.

[0056] In summary, by executing the multi-source vulnerability data fusion processing method in steps 201 to 204 in the embodiments of the present disclosure, after receiving multiple vulnerability data discovered through different channels, each vulnerability data can be separately fused. During the process of processing the current first vulnerability data, if the first vulnerability data carries a corresponding first knowledge base identifier, since the embodiments of the present disclosure have pre-divided the hierarchical order of multiple knowledge base identifiers and defined that the level of the first type of identifier is higher than that of the second type of identifier, that is, the vulnerability knowledge base level corresponding to the first type of identifier is higher than the vulnerability knowledge base corresponding to the second type of identifier. During the process of sequentially matching the first knowledge base identifier from high to low according to the hierarchical order of multiple knowledge base identifiers, it is possible to first determine whether the first vulnerability data belongs to the first type of identifier, and then determine whether it belongs to the second type of identifier after it does not belong to the first type of identifier. Then, after the first vulnerability data belongs to the second type of identifier, it is determined whether there is a first type of identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge base identifier, so as to determine whether the first vulnerability data is the same as the vulnerability data indicated by the first type of identifier. Once there is a first type of identifier with a mapping relationship, it means that although the first vulnerability data is scanned through the lower-level vulnerability knowledge base corresponding to the second type of identifier, it is the same as the vulnerability data marked by the higher-level vulnerability knowledge base corresponding to the first type of identifier. Then, during the fusion process, the first type of identifier with a mapping relationship can be used as the target knowledge base identifier to allocate the first vulnerability data under the higher-level target knowledge base identifier to obtain the corresponding vulnerability fusion result. Compared with the solutions in the related art, the embodiments of the present disclosure can divide the levels of different vulnerability knowledge bases, and when there are differences in the vulnerability data scanned through different channels, identify whether the vulnerability data scanned at a lower level is the same as the vulnerability data under the higher-level identifier, so as to effectively integrate multi-source vulnerability data, reduce the probability of false reporting or missing reporting of vulnerability data caused by significant differences in the scanning results of different channels, and ultimately improve network security.

[0057] Next, the further included content in steps 201 to 204 in the embodiments of the present disclosure will be described in detail.

[0058] Please refer to Figure 4 , Figure 4 is Figure 2 a schematic flowchart of the further included process after step 202 in. In some embodiments, after sequentially performing fusion processing on each vulnerability data and for the current first vulnerability data being fused, the multi-source vulnerability data fusion processing method may include steps 301 to 303: Step 301, if the first vulnerability data does not carry a corresponding first knowledge base identifier, obtain multiple second vulnerability data from the preset vulnerability fusion table; Among them, the second knowledge base identifier of the second vulnerability data does not belong to the first type of identifier nor the second type of identifier; Step 302: Determine the vulnerability similarity between the first vulnerability data and each second vulnerability data to obtain the corresponding similarity value; Step 303: When any one of the similarity values is greater than or equal to the preset first similarity threshold, determine the second knowledge base identifier of the similar second vulnerability data as the target knowledge base identifier, and assign the first vulnerability data under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0059] In the above steps, the second vulnerability data is the vulnerability data obtained from the preset vulnerability fusion table. Its feature is that the second knowledge base identifier corresponding to this vulnerability data belongs to neither the first type of identifier nor the second type of identifier. Like the first vulnerability data being processed currently, it is also a kind of vulnerability data and is used for subsequent comparison and analysis. The second knowledge base identifier is the corresponding knowledge base identifier carried by the second vulnerability data, which is used to indicate which specific vulnerability knowledge base scanned this vulnerability data.

[0060] It should be noted that if the first vulnerability data does not carry the corresponding first knowledge base identifier, it indicates that this vulnerability data is not scanned based on the vulnerability knowledge base through any means. Then this first vulnerability data is a non-standard vulnerability. At this time, in order to more comprehensively fuse this vulnerability data, it is necessary to refer to other vulnerability data that has been fused and whose knowledge base identifier also does not belong to these two types. By comparing their similarities, the reasonable attribution of the first vulnerability data can be determined to more accurately integrate the vulnerability data.

[0061] The vulnerability similarity is an index used to measure the similarity degree between the first vulnerability data and each second vulnerability data. By calculating the vulnerability similarity, the similarity of these vulnerability data in multiple aspects can be judged, so as to determine whether they belong to the same category or the same vulnerability. The similarity value is the specific value representing the vulnerability similarity. It is calculated through a specific algorithm or method. The larger the value, the higher the similarity degree between the two vulnerability data.

[0062] The preset first similarity threshold is a preset value, which is used as the standard for judging whether the first vulnerability data and the second vulnerability data are similar enough. If the calculated similarity value is greater than or equal to this threshold, it is considered that the two vulnerability data have a high similarity and can be classified into one category.

[0063] It should be noted that after obtaining the similarity values between the first vulnerability data and each second vulnerability data in the embodiments of the present disclosure, a standard is required to determine whether these similarities are high enough to decide whether to classify the first vulnerability data and a certain second vulnerability data into the same category. The preset first similarity threshold is such a standard. When a certain similarity value is greater than or equal to the threshold, it indicates that the first vulnerability data and the corresponding second vulnerability data are very similar, and they are very likely to be the same type or the same vulnerability. Therefore, the second knowledge base identifier of the similar second vulnerability data is determined as the target knowledge base identifier, and the first vulnerability data is assigned under this identifier. On the contrary, when any similarity value is less than the first similarity threshold, it indicates that there is no second vulnerability data similar to the first vulnerability data. Then, the embodiments of the present disclosure can configure a corresponding custom identifier for the first vulnerability data to distinguish it from other vulnerability data in the vulnerability fusion table, and use this custom identifier as the target knowledge base identifier and assign the first vulnerability data under this identifier. Based on this, further effective integration of vulnerability data can be achieved, improving the accuracy and rationality of vulnerability data management.

[0064] For example, taking the case where the vulnerability data is scanned based on CVE, CNNVD, and CNVD as an example, the CVE ID, CNNVD ID, and CNVD ID constitute a multi-level knowledge base identifier. Among them, multiple identifiers under CVE are defined as the first type of identifier, while multiple identifiers under CNNVD and CNVD are defined as the second type of identifier. When the first knowledge base identifier does not match any CVE ID, and the first knowledge base identifier also does not match any CNNVD ID and CNVD ID, it indicates that this vulnerability is a non-standard vulnerability identified by other means, such as non-standard vulnerabilities determined through penetration testing, static analysis, dynamic analysis, etc. Therefore, the embodiments of the present disclosure calculate the similarity value between this current non-standard vulnerability and other non-standard vulnerabilities (i.e., non-CVE, non-CNNVD, non-CNVD records) in the vulnerability fusion table. Subsequently, based on whether this similarity value reaches or exceeds the preset threshold, it is decided whether to add this vulnerability to the vulnerability fusion table or update and modify relevant fields such as the vulnerability scanning tool of the corresponding vulnerability record in the vulnerability fusion table.

[0065] Therefore, the embodiments of the present disclosure can systematically process the scanning results from different sources, especially different vulnerability scanning tools, and achieve effective fusion of scanning results from different sources in the same network environment. This fusion process not only integrates multiple vulnerability information sources but also significantly reduces the false negative risk caused by differences in sources, thus providing a more comprehensive and accurate basis for the assessment of the network security status.

[0066] Please refer to Figure 5 , Figure 5 is Figure 4The process schematic diagram further included in step 302. In some embodiments, in the process of determining the vulnerability similarity between the first vulnerability data and each second vulnerability data to obtain the corresponding similarity value, steps 401 to 403 may be included: Step 401: respectively determine the vulnerability attributes of multiple different attribute types under the first vulnerability data and the second vulnerability data, and configure corresponding similarity algorithms for different attribute types; Step 402: for each of the first vulnerability data and each second vulnerability data, under the corresponding similarity algorithm, sequentially determine the attribute similarity between the vulnerability attributes under the same attribute type to obtain sub-similarity values under multiple attribute types; Step 403: based on the weights under each attribute type, perform weighted calculation on multiple sub-similarity values to obtain the similarity value between the first vulnerability data and each second vulnerability data.

[0067] In the above steps, the vulnerability attribute is various characteristic information included in the vulnerability data, and different vulnerability attributes can describe the characteristics of the vulnerability from different perspectives. The attribute type is the classification of the vulnerability attributes. For example, according to the technical characteristics of the vulnerability, it can be divided into attribute types such as code execution type and information leakage type; according to the impact degree of the vulnerability, it can be divided into attribute types such as severe, medium, and minor; according to the types of attributes, it can be divided into including component name, component version range, vulnerability description, repair suggestion, vulnerability type, exploitation method, and impact range, etc. The similarity algorithm is a method or formula for calculating the similarity degree between two vulnerability data under a certain specific attribute type. Different attribute types may require different similarity algorithms.

[0068] It should be noted that in order to accurately calculate the similarity between the first vulnerability data and each second vulnerability data, they need to be compared from multiple aspects. By determining the vulnerability attributes of different attribute types, the characteristics of the vulnerability can be comprehensively understood. Because different attribute types have different characteristics and measurement criteria, using a unified algorithm may not accurately reflect the similarity degree between vulnerabilities, and configuring corresponding similarity algorithms for different attribute types can select the most appropriate calculation method according to the characteristics of the attributes, thereby improving the accuracy and reliability of the similarity calculation.

[0069] The attribute similarity is a numerical representation of the degree of similarity between the corresponding vulnerability attributes of the first vulnerability data and the second vulnerability data under a specific attribute type. For example, under the attribute type of vulnerability type, the attribute similarity between two vulnerabilities is the result obtained by determining whether their vulnerability types are the same. The sub-similarity value refers to the specific numerical value of the attribute similarity calculated under each attribute type. By sequentially calculating the attribute similarities between the vulnerability attributes under the same attribute type, multiple sub-similarity values can be obtained. These sub-similarity values reflect the similarity situation of the two vulnerabilities from different aspects and provide basic data for subsequent comprehensive calculation of the overall similarity.

[0070] The weight is a numerical value set for each attribute type and is used to represent the importance of the attribute type in the comprehensive calculation of similarity. Different attribute types may have different importance in judging the similarity of vulnerabilities. For example, in some cases, the vulnerability type may be more important than the discovery time, so the weight of the vulnerability type can be set higher.

[0071] It should be noted that since different attribute types have different importance in judging the vulnerability similarity, simply adding or averaging the individual sub-similarity values may not accurately reflect the actual similarity degree of the two vulnerabilities. In the embodiments of the present disclosure, based on the weights of each attribute type, multiple sub-similarity values are weighted and calculated to obtain the similarity value between the first vulnerability data and each second vulnerability data. In this way, by setting weights for each attribute type and performing weighted calculation on multiple sub-similarity values, they can be reasonably integrated according to the importance of the attributes, so as to obtain a more accurate similarity value that can better reflect the true similarity between vulnerabilities, providing a more reliable basis for subsequent determination of the attribution of vulnerability data.

[0072] In the embodiments of the present disclosure, taking the attribute types of component name, component version, vulnerability description, repair suggestion, vulnerability type, exploitation method, and affected scope as examples, the component name refers to the specific name of the software component with a vulnerability, which is used to clearly point out the specific component where the problem lies, such as a specific software library, module, or application, etc.; the component version is the version number of the component with a vulnerability. Components of different versions may have different characteristics and code logics. Clearly stating the version number helps to accurately locate and distinguish the problems existing in different versions; the vulnerability description details the specific situation of the vulnerability, including how the vulnerability was discovered, the manifestation form of the vulnerability, and the possible consequences, etc., so that relevant personnel can clearly understand the nature and harm of the vulnerability; the repair suggestion is the specific solution and measure proposed for the vulnerability, guiding technicians on how to repair the vulnerability to eliminate security risks; the vulnerability type classifies the vulnerability. For example, common types include injection vulnerabilities, out-of-bounds access vulnerabilities, code execution vulnerabilities, etc. This helps to understand the nature and characteristics of the vulnerability from a macro perspective, so as to adopt corresponding prevention and repair strategies; the exploitation method describes the specific methods and ways that malicious attackers may use the vulnerability. Understanding the exploitation method can help security personnel better assess risks and take targeted protection measures; the affected scope clarifies the scope of systems, applications, business functions, etc. that the vulnerability may affect, helping relevant personnel comprehensively evaluate the harm degree of the vulnerability, so as to determine the repair priority and take corresponding emergency measures.

[0073] Please refer to Figure 6 , Figure 6 is Figure 5 the process schematic diagram further included in step 402 in. In some embodiments, in the process of sequentially determining the attribute similarity between vulnerability attributes under the same attribute type to obtain sub-similarity values under multiple attribute types, steps 501 to 507 may be included: Step 501, calculate the edit distance and the first Jaccard similarity value between vulnerability attributes with the attribute type of component name, and determine the sub-similarity value under the component name by combining the edit distance and the first Jaccard similarity value; Step 502, calculate the hierarchical version split vector and the first cosine similarity value between vulnerability attributes with the attribute type of component version, and determine the sub-similarity value under the component version by combining the hierarchical version split vector and the first cosine similarity value; Step 503, calculate the first domain-enhanced semantic vector feature and the second cosine similarity value between vulnerability attributes with the attribute type of vulnerability description, and determine the sub-similarity value under the vulnerability description by combining the first domain-enhanced semantic vector feature and the second cosine similarity value; Step 504: Calculate the action and object semantic role annotation results and the third cosine similarity value among vulnerability attributes with the property type of repair suggestions, and determine the sub-similarity value under the repair suggestions by combining the annotation results and the third cosine similarity value; Step 505: Calculate the classification tree path distance or the second domain-enhanced semantic vector feature among vulnerability attributes with the property type of vulnerability types, and determine the sub-similarity value under the vulnerability types by combining the classification tree path distance or the second domain-enhanced semantic vector feature; Step 506: Calculate the TF-IDF vector feature and the fourth cosine similarity value among vulnerability attributes with the property type of exploitation methods, and determine the sub-similarity value under the exploitation methods by combining the TF-IDF vector feature and the fourth cosine similarity value; Step 507: Calculate the second Jaccard similarity value of the asset types among vulnerability attributes with the property type of impact scope, and determine the sub-similarity value under the impact scope based on the second Jaccard similarity value.

[0074] In the above steps, the edit distance refers to the minimum number of edit operations (insertion, deletion, replacement) required to convert one string to another. When comparing component names, the edit distance can measure the degree of difference between two component names. The first Jaccard similarity value is used to measure the similarity degree of two sets, and the calculation formula is the number of intersection elements of the two sets divided by the number of union elements. For component names, it can be split into character sets to calculate the Jaccard similarity. It should be noted that the component name is an important attribute of the vulnerability, and a single measurement method may not accurately reflect its similarity. Therefore, the edit distance focuses on the differences at the character level, and the Jaccard similarity focuses on the overlap of set elements. Combining the two can more comprehensively and accurately determine the similarity degree of component names.

[0075] The hierarchical version split vector splits the component version number into a vector form by levels (such as major version number, minor version number, revision number, etc.) to facilitate mathematical calculations and comparisons. The first cosine similarity value is used to measure the cosine value of the angle between two vectors. The closer the value is to 1, the more similar the two vectors are. When comparing component versions, the cosine similarity of the hierarchical version split vector is calculated to judge the similarity degree of the versions. It should be noted that the similarity of component versions is very important for judging whether vulnerabilities are the same or related. The hierarchical version split vector can clearly represent the information of each level of the version, and the cosine similarity can quantify the similarity degree between versions. Combining the two can more accurately determine the sub-similarity value of component versions.

[0076] The semantic vector features of the first domain enhancement (SBERT) are in the field of network security. For the vulnerability description, semantic analysis is performed, it is converted into vector form, and enhanced by combining domain knowledge to better represent the semantic information of the vulnerability description. The second cosine similarity value is also the cosine value of the included angle between two vectors, which is used here to compare the similarity degree of the semantic vector features of the first domain enhancement of two vulnerability descriptions. It should be noted that the vulnerability description usually contains detailed information, and it is difficult for ordinary text matching to accurately judge its similarity. Through the domain-enhanced semantic vector features, the semantic information of the vulnerability description can be captured, and the cosine similarity can be combined to more accurately measure the similarity degree of two vulnerability descriptions.

[0077] The semantic role annotation results of actions and objects are to perform semantic role annotation on the actions (such as update, delete, etc.) and objects (such as components, files, etc.) in the repair suggestions to clarify the semantic roles of each component in the sentence. The third cosine similarity value is used to compare the similarity degree of the semantic role annotation result vectors of two repair suggestions. It should be noted that the similarity of the repair suggestions can reflect the similarity of the vulnerabilities. Through the semantic role annotation of actions and objects, the key semantic information of the repair suggestions can be extracted, and the cosine similarity can be combined to more accurately judge the similarity degree of the repair suggestions.

[0078] The path distance of the classification tree (such as the CWE classification tree) is the shortest path length between two vulnerability type nodes in the classification tree of vulnerability types. The shorter the path distance (the shortest common ancestor distance), the more similar the two vulnerability types are. The second domain-enhanced semantic vector features are similar to the first domain-enhanced semantic vector features and are the vector representations after semantic analysis of the vulnerability types and enhancement by combining domain knowledge. It should be noted that the vulnerability type is an important basis for judging the similarity of vulnerabilities. The classification tree path distance measures the similarity of vulnerability types from the structure, and the second domain-enhanced semantic vector features measure the similarity from the semantics. Combining the two can more comprehensively and accurately determine the sub-similarity value of the vulnerability types.

[0079] TF-IDF (Term Frequency-Inverse Document Frequency) is a commonly used weighting technique for information retrieval and text mining. By converting the text of the exploitation method into TF-IDF vector features to represent the importance of each word in the text. The fourth cosine similarity value is used to compare the similarity degree of the TF-IDF vector features of two exploitation methods. It should be noted that the exploitation method describes how an attacker exploits a vulnerability, and its similarity is very important for judging the relevance of vulnerabilities. The TF-IDF vector features can highlight the important words in the exploitation method text, and the cosine similarity can be combined to more accurately judge the similarity degree of the exploitation methods.

[0080] The second Jaccard similarity value also measures the similarity between two sets and is used here to compare the sets of asset types within two scopes of influence. It should be noted that the asset types within the scope of influence are an important aspect in judging vulnerability similarity. By calculating the Jaccard similarity of asset types, the similarity between two scopes of influence can be directly measured, thereby determining the sub-similarity value under the scope of influence.

[0081] When calculating the sub-similarity values for each attribute type as described above, all sub-similarity values need to be mapped to the interval from 0 to 1. And in the embodiments of the present disclosure, by calculating the sub-similarity values for these attribute types, the similarity between vulnerability data is comprehensively evaluated from multiple dimensions, providing a detailed and accurate basis for the fusion of multi-source vulnerability data, helping to improve the integration effect of vulnerability data, reducing the probability of false positives and false negatives, and enhancing the accuracy of network security assessment.

[0082] Furthermore, in combination with the importance of relevant attributes, corresponding weights are configured for different vulnerability attribute types, and then the overall similarity is calculated through weighted averaging to obtain the similarity values between pairs of vulnerability data. The similarity value calculation formula is: Similarity = 0.20 * Sim 组件名称 + 0.15 * Sim 组件版本 + 0.20 * Sim 漏洞描述 + 0.15 * Sim 修复建议 + 0.10 * Sim 漏洞类型 + 0.10 * Sim 利用方式 + 0.10 * Sim 影响范围 , where each Sim value is the sub-similarity value calculated for each attribute type in the above embodiments.

[0083] Through the above method, the similarity value between the current first vulnerability data and the vulnerability data of the non-standard record in the vulnerability fusion table, that is, the second vulnerability data, is calculated. This is also called the overall similarity between the two vulnerability data. If the overall similarity is greater than or equal to a set threshold, such as greater than or equal to 0.75, then the second knowledge base identifier of the similar second vulnerability data is determined as the target knowledge base identifier, and the first vulnerability data is assigned under the target knowledge base identifier to obtain the corresponding vulnerability fusion result; otherwise, if the overall similarity is less than the threshold, such as less than 0.75, then according to the vulnerability details, a new vulnerability record is created in accordance with the unified vulnerability description format and saved in the vulnerability fusion table.

[0084] In some embodiments, after determining whether there is a first type of identifier having a mapping relationship with the first knowledge base identifier in a preset mapping knowledge base, the multi-source vulnerability data fusion processing method may further include step 601: Step 601, when there is no first - type identifier in the mapping knowledge base that has a mapping relationship with the first knowledge - base identifier, use the second - type identifier that matches the first knowledge - base identifier as the target knowledge - base identifier, and allocate the first vulnerability data under the target knowledge - base identifier to obtain the corresponding vulnerability fusion result.

[0085] It should be noted that during the multi - source vulnerability data fusion process, when matching the first knowledge - base identifier of the first vulnerability data, it is found that it does not match any of the first - type identifiers, and there is also no first - type identifier in the mapping knowledge base that has a mapping relationship with the first knowledge - base identifier. This indicates that the first vulnerability data being processed is scanned from a vulnerability knowledge base corresponding to a second - type identifier of a lower level, and this vulnerability data has no association with the vulnerability data represented by the first - type identifier of a higher level. In this case, in order to reasonably classify and manage the vulnerability data, it is necessary to allocate it under a suitable knowledge - base identifier.

[0086] Since it cannot be classified under the first - type identifier of a higher level at this time, the embodiment of the present disclosure uses the second - type identifier that matches the first knowledge - base identifier as the target knowledge - base identifier and allocates the first vulnerability data under this target knowledge - base identifier. In this way, the vulnerability data can be effectively integrated, ensuring that each vulnerability data has a reasonable attribution, facilitating subsequent network security analysis and processing, and improving the usability and management efficiency of the vulnerability data.

[0087] Furthermore, the multiple second - type identifiers in the embodiment of the present disclosure include multiple second - type identifiers under the first vulnerability knowledge base and multiple second - type identifiers under the second vulnerability knowledge base at the same level as the first vulnerability knowledge base. Both the first vulnerability knowledge base and the second vulnerability knowledge base belong to low - level vulnerability knowledge bases. In this embodiment, when the high - level vulnerability knowledge base is CVE, the first vulnerability knowledge base is one of CNNVD and CNVD, and the second vulnerability knowledge base is the other one of CNNVD and CNVD. The division of the first vulnerability knowledge base and the second vulnerability knowledge base can be carried out according to actual needs, and the embodiment of the present disclosure does not make specific limitations in this regard.

[0088] Please refer to Figure 7 , Figure 7 which is a schematic flowchart further included after step 601 in the embodiment of the present disclosure. In some embodiments, after determining whether there is a first - type identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge - base identifier, the multi - source vulnerability data fusion processing method may further include steps 701 to 702: Step 701: When the first knowledge base identifier matches one of the second - type identifiers under a first vulnerability knowledge base, determine whether there is a second - type identifier under a second vulnerability knowledge base that has a mapping relationship with the first knowledge base identifier in the mapping knowledge base; Step 702: When there is a second - type identifier under a second vulnerability knowledge base that has a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the second - type identifier under the second vulnerability knowledge base with the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data to the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0089] In the above steps, in the multi - source vulnerability data fusion process, when it has been determined that the first knowledge base identifier does not match any first - type identifier, and there is also no first - type identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, at this time, the first vulnerability data has been allocated to the second - type identifier (belonging to the low - level knowledge base identifier) that matches the first knowledge base identifier. However, since there are multiple low - level vulnerability knowledge bases (such as CNNVD and CNVD), there may be some situations where the vulnerability data between these low - level knowledge bases is essentially the same, but only has different identifiers because of the different knowledge bases it depends on. Therefore, in order to further accurately integrate the vulnerability data, it is necessary to determine whether the second - type identifier under the first vulnerability knowledge base that the first vulnerability data matches has a mapping relationship with the second - type identifier under the second vulnerability knowledge base of the same level, so as to classify the same vulnerability data more reasonably, improve the integration effect of the vulnerability data, and avoid false alarms or missed reports of vulnerability data caused by differences between low - level knowledge bases.

[0090] When it is found in the mapping knowledge base that there is a second - type identifier under a second vulnerability knowledge base that has a mapping relationship with the first knowledge base identifier, it means that the current first vulnerability data is essentially the same vulnerability as a certain vulnerability data scanned from another low - level vulnerability knowledge base (the second vulnerability knowledge base). In order to make the classification of the vulnerability data more reasonable and accurate, use the second - type identifier under the second vulnerability knowledge base with the mapping relationship as the target knowledge base identifier, and re - allocate the first vulnerability data to this new target knowledge base identifier. In this way, the same vulnerability data can be integrated under the same low - level knowledge base identifier, which is convenient for subsequent management, analysis, and processing of the vulnerability data, reduces the confusion caused by differences between low - level knowledge bases, and improves the accuracy of network security assessment.

[0091] For example, vulnerability scanning tool X performs vulnerability scanning based on CNNVD (the first vulnerability knowledge base) and discovers vulnerability x. The first knowledge base identifier of this vulnerability data is a certain CNNVD ID. When performing fusion processing on vulnerability x, it has been determined that this CNNVD ID does not match any CVE ID, and there is no CVE ID with a mapping relationship to this CNNVD ID in the mapping knowledge base. Then, step 701 is executed. At this time, it is necessary to check in the mapping knowledge base whether there is a second type of identifier (CNVD ID) under CNVD that has a mapping relationship with this CNNVD ID. If a CNVD ID with a mapping relationship to the CNNVD ID of vulnerability x is found in the mapping knowledge base, then this CNVD ID is used as the target knowledge base identifier, and vulnerability x is assigned under this CNVD ID to obtain the corresponding vulnerability fusion result. That is to say, originally, vulnerability x assigned under the CNNVD ID is reclassified under the CNVD ID because it is found that it has the same vulnerability essence as the vulnerability represented by a certain CNVD ID, thus completing a more accurate fusion processing of vulnerability x.

[0092] Please refer to Figure 8 , Figure 8 is Figure 2 Another process schematic diagram further included after step 202 in []. In some embodiments, after sequentially matching the first knowledge base identifiers in the order of the levels of the pre-arranged multi-level knowledge base identifiers from high to low, the method for fusing multi-source vulnerability data may include step 801: Step 801, when the first knowledge base identifier matches one of the first type of identifiers, use the first type of identifier that matches the first knowledge base identifier as the target knowledge base identifier, and assign the first vulnerability data under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

[0093] In the above step, when the first knowledge base identifier can match the first type of identifier, it indicates that the vulnerability data comes from a high-level vulnerability knowledge base. Using the first type of identifier that matches it as the target knowledge base identifier and assigning the vulnerability data can more accurately classify and integrate the vulnerability data, improve the accuracy and reliability of vulnerability data classification, and provide a more reliable basis for subsequent network security analysis and decision-making.

[0094] Please refer to Figure 9 , Figure 9 is Figure 2 The process schematic diagram further included in step 204 in []. In some embodiments, in the process of assigning the first vulnerability data under the target knowledge base identifier to obtain the corresponding vulnerability fusion result, steps 901 to 903 may be included: Step 901, determine whether there is a vulnerability storage record of the target knowledge base identifier in the preset vulnerability fusion table; Step 902, when there is a vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table, update the vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table based on the first vulnerability data to obtain the corresponding vulnerability fusion result; Step 903, when there is no vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table, obtain the vulnerability description information of the first vulnerability data from the vulnerability knowledge base corresponding to the target knowledge base identifier, and create a vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table based on the vulnerability description information to obtain the corresponding vulnerability fusion result.

[0095] In the above steps, the vulnerability storage record refers to the relevant information about vulnerabilities recorded in the vulnerability fusion table for a certain target knowledge base identifier, including but not limited to the vulnerability data under the target knowledge base identifier and the vulnerability scanning tool that scanned the vulnerability data, etc. When allocating the first vulnerability data under the target knowledge base identifier, it is first necessary to determine whether there is already a vulnerability storage record about the target knowledge base identifier in the preset vulnerability fusion table. This is because if there is already a record, it means that the vulnerability data related to the target knowledge base identifier has been processed before, and the subsequent processing method will be different from the case where there is no record. Through this step of judgment, the processing strategy for the first vulnerability data can be determined, so as to update the vulnerability fusion table more accurately and obtain the correct vulnerability fusion result.

[0096] When there is a vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table, it means that the vulnerability data related to the target knowledge base identifier has been processed and recorded in the table before. At this time, updating the vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table based on the first vulnerability data is to supplement the new vulnerability data information into the existing record, making the information in the vulnerability fusion table more complete and accurate. This can ensure that all relevant vulnerability data under the same target knowledge base identifier can be reflected in the vulnerability fusion table, providing more comprehensive data support for subsequent network security analysis.

[0097] The vulnerability description information refers to the detailed information about the first vulnerability data obtained from the vulnerability knowledge base corresponding to the target knowledge base identifier, including the detailed description of the vulnerability, the affected scope, the repair suggestions, etc. These information are very important for accurately creating a vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table.

[0098] When there is no vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table, it is necessary to obtain the vulnerability description information of the first vulnerability data from the vulnerability knowledge base corresponding to the target knowledge base identifier, and create a vulnerability storage record with the target knowledge base identifier in the vulnerability fusion table based on this information. This is because if there is no relevant record, a new record item needs to be established in the vulnerability fusion table for this new target knowledge base identifier and its corresponding first vulnerability data, so as to manage and analyze it later. By obtaining the vulnerability description information, the newly created record can be made more complete and accurate, providing basic data for subsequent network security analysis.

[0099] Please refer to Figure 10 , Figure 10 which is another process schematic diagram of the multi-source vulnerability data fusion processing method provided by the embodiments of the present disclosure. In some embodiments, the multi-source vulnerability data fusion processing method may further include steps 1001 to 1002: Step 1001, determine the vulnerability similarity between the vulnerability data under any one first type of identifier and the vulnerability data under any one second type of identifier, and obtain the corresponding similarity value; Step 1002, when any one similarity value is greater than or equal to a preset second similarity threshold, establish a mapping relationship between the corresponding first type of identifier and the second type of identifier for the determined similar vulnerability data, and save it to the mapping knowledge base.

[0100] In the above steps, the vulnerability similarity is an index used to measure the similarity degree between the vulnerability data under any one first type of identifier and the vulnerability data under any one second type of identifier. It reflects the similarity of the two vulnerability data in terms of vulnerability characteristics, influence range, harm degree, etc. The similarity value in this step is a specific value obtained by calculating the vulnerability similarity. This value can intuitively represent the similarity degree between the two vulnerability data. The larger the value, the higher the similarity. It should be noted that the method for calculating the similarity value in this step can refer to the method for calculating the similarity value based on multiple different attribute types in the above embodiments, and will not be elaborated here.

[0101] In the multi-source vulnerability data fusion processing, the vulnerability data scanned by knowledge bases of different levels may be similar or even the same. Determining the vulnerability similarity between the vulnerability data under any one first type of identifier and the vulnerability data under any one second type of identifier, and obtaining the corresponding similarity value, is to find out those vulnerability data that may be essentially the same or similar although seemingly from knowledge bases of different levels.

[0102] The preset second similarity threshold is a numerically predefined standard for determining whether the similarity of vulnerability data is high enough. When the calculated similarity value is greater than or equal to this threshold, it is considered that two pieces of vulnerability data are similar.

[0103] When any similarity value is greater than or equal to the preset second similarity threshold, it indicates that the vulnerability data under the corresponding first type of identifier and the vulnerability data under the second type of identifier have a high degree of similarity and are very likely to be the same or similar vulnerabilities in essence, but only have different identifiers due to relying on different knowledge bases; conversely, if the similarity value is less than the preset second similarity threshold, it indicates that the two pieces of vulnerability data being compared are not similar, so the vulnerability data being compared is not processed. In the embodiments of the present disclosure, a mapping relationship between the corresponding first type of identifier and the second type of identifier is established for the determined similar vulnerability data and saved in the mapping knowledge base. Doing so can further improve the content of the mapping knowledge base.

[0104] Based on this, the mapping knowledge base in the embodiments of the present disclosure records the relationships between different knowledge base identifiers. Improving it helps to more accurately determine whether the vulnerability data scanned by a low-level knowledge base is related to the vulnerability data in a high-level knowledge base in subsequent vulnerability data fusion processing, so as to more effectively integrate multi-source vulnerability data and reduce the probability of false positives and false negatives.

[0105] Please refer to Figure 11 , the embodiments of the present disclosure also provide a multi-source vulnerability data fusion processing device, which can implement the above multi-source vulnerability data fusion processing method. The multi-source vulnerability data fusion processing device includes: A multi-source data acquisition module 1101, configured to respectively acquire multiple pieces of vulnerability data discovered through multiple different channels. Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; An identifier matching module 1102, configured to perform fusion processing on each piece of vulnerability data in sequence, and for the first piece of vulnerability data currently undergoing fusion processing, if the first piece of vulnerability data carries the corresponding first knowledge base identifier, match the first knowledge base identifier in descending order according to the level order of the pre-arranged multi-level knowledge base identifiers. Among them, the multi-level knowledge base identifiers include multiple first types of identifiers and multiple second types of identifiers, and the level of the first type of identifier is higher than that of the second type of identifier; A mapping determination module 1103, configured to determine whether there is a first type of identifier having a mapping relationship with the first knowledge base identifier in the preset mapping knowledge base when the first knowledge base identifier does not match any first type of identifier but matches one of the second type of identifiers; A result determination module 1104, configured to, when there is a first type of identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, use the first type of identifier with the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

[0106] In summary, the multi-source vulnerability data fusion processing device, by executing the multi-source vulnerability data fusion processing method in the above embodiments, can perform fusion processing on each vulnerability data respectively after receiving multiple vulnerability data discovered through different channels. During the process of processing the current first vulnerability data, if the first vulnerability data carries a corresponding first knowledge base identifier, since the embodiments of the present disclosure have pre-divided the hierarchical order of multiple knowledge base identifiers and limited that the level of the first type of identifier is higher than that of the second type of identifier, that is, the vulnerability knowledge base level corresponding to the first type of identifier is higher than the vulnerability knowledge base corresponding to the second type of identifier. During the process of sequentially matching the first knowledge base identifier from high to low according to the hierarchical order of multiple knowledge base identifiers, it is possible to first determine whether the first vulnerability data belongs to the first type of identifier, and then determine whether it belongs to the second type of identifier after it does not belong to the first type of identifier. Then, after the first vulnerability data belongs to the second type of identifier, it is determined whether there is a first type of identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge base identifier, so as to determine whether the first vulnerability data is the same as the vulnerability data indicated by the first type of identifier. Once there is a first type of identifier with a mapping relationship, it means that although the first vulnerability data is scanned through the lower-level vulnerability knowledge base corresponding to the second type of identifier, it is the same as the vulnerability data marked by the higher-level vulnerability knowledge base corresponding to the first type of identifier. Then, during the fusion process, the first type of identifier with the mapping relationship can be used as the target knowledge base identifier to achieve allocating the first vulnerability data under the higher-level target knowledge base identifier to obtain a corresponding vulnerability fusion result. Compared with the solutions in the related art, the embodiments of the present disclosure can, by dividing the levels of different vulnerability knowledge bases, identify whether the vulnerability data scanned at a lower level is the same as the vulnerability data under a higher-level identifier when there are differences in the vulnerability data scanned through different channels, so as to effectively integrate multi-source vulnerability data, reduce the probability of false alarms or missed reports of vulnerability data caused by significant differences in scanning results through different channels, and ultimately improve network security.

[0107] The specific implementation manner of the multi-source vulnerability data fusion processing device is basically the same as the specific embodiments of the above multi-source vulnerability data fusion processing method, and will not be elaborated here. On the premise of meeting the requirements of the embodiments of the present disclosure, other functional modules can also be set in the multi-source vulnerability data fusion processing device to implement the multi-source vulnerability data fusion processing method in the above embodiments.

[0108] Embodiments of the present disclosure also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned multi-source vulnerability data fusion processing method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0109] Please refer to Figure 12 , Figure 12 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes: A processor 1201, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; A memory 1202, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1202 can store an operating device and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute the multi-source vulnerability data fusion processing method of the embodiments of the present disclosure; An input / output interface 1203, which is used to implement information input and output; A communication interface 1204, which is used to implement communication interaction between this device and other devices. Communication can be achieved through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); A bus 1205, which transmits information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204); Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 achieve communication connections with each other inside the device through the bus 1205.

[0110] Embodiments of the present disclosure also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned multi-source vulnerability data fusion processing method is implemented.

[0111] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0112] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0113] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the devices and equipment, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0116] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above figures are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or equipment.

[0117] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0118] In several embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0119] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0120] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0121] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0122] The preferred embodiments of the embodiments of the present disclosure have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present disclosure shall be within the scope of rights of the embodiments of the present disclosure.

Claims

1. A method for fusing and processing multi-source vulnerability data, characterized in that, Including: Obtain multiple pieces of vulnerability data discovered through multiple different channels respectively. Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; Perform fusion processing on each piece of the vulnerability data in sequence. For the first piece of vulnerability data undergoing fusion processing currently, if the first piece of vulnerability data carries a corresponding first knowledge base identifier, match the first knowledge base identifier in descending order according to the pre-arranged hierarchical order of multiple levels of knowledge base identifiers. Among them, the multiple levels of knowledge base identifiers include multiple first-type identifiers and multiple second-type identifiers, and the level of the first-type identifier is higher than that of the second-type identifier; When the first knowledge base identifier does not match any of the first-type identifiers but matches one of the second-type identifiers, determine whether there is a first-type identifier having a mapping relationship with the first knowledge base identifier in a preset mapping knowledge base; When there is a first-type identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the first-type identifier having the mapping relationship as the target knowledge base identifier, and allocate the first piece of vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

2. The fusion processing method for multi-source vulnerability data according to claim 1, characterized in that After performing the fusion processing on each piece of the vulnerability data in sequence and for the first piece of vulnerability data undergoing fusion processing currently, the multi-source vulnerability data fusion processing method further includes: If the first piece of vulnerability data does not carry a corresponding first knowledge base identifier, obtain multiple second pieces of vulnerability data from a preset vulnerability fusion table. Among them, the second knowledge base identifier of the second piece of vulnerability data does not belong to the first-type identifier nor the second-type identifier; Determine the vulnerability similarity between the first piece of vulnerability data and each second piece of vulnerability data to obtain corresponding similarity values; When any one of the similarity values is greater than or equal to a preset first similarity threshold, determine the second knowledge base identifier of the similar second piece of vulnerability data as the target knowledge base identifier, and allocate the first piece of vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

3. The fusion processing method for multi-source vulnerability data according to claim 2, wherein The determining the vulnerability similarity between the first piece of vulnerability data and each second piece of vulnerability data to obtain corresponding similarity values includes: Determine the vulnerability attributes of multiple different attribute types under the first piece of vulnerability data and the second piece of vulnerability data respectively, and configure corresponding similarity algorithms for different attribute types; For each of the first piece of vulnerability data and each second piece of vulnerability data, under the corresponding similarity algorithm, determine the attribute similarity between the vulnerability attributes under the same attribute type in sequence to obtain sub-similarity values under multiple attribute types; Based on the weights under each attribute type, perform weighted calculation on multiple sub-similarity values to obtain the similarity value between the first piece of vulnerability data and each second piece of vulnerability data.

4. The fusion processing method for multi-source vulnerability data according to claim 3, wherein Sequentially determining the attribute similarity between the vulnerability attributes under the same attribute type to obtain sub-similarity values under multiple attribute types, including: Calculating the edit distance and the first Jaccard similarity value between the vulnerability attributes with the attribute type of component name, and determining the sub-similarity value under the component name by combining the edit distance and the first Jaccard similarity value; Calculating the hierarchical version splitting vector and the first cosine similarity value between the vulnerability attributes with the attribute type of component version, and determining the sub-similarity value under the component version by combining the hierarchical version splitting vector and the first cosine similarity value; Calculating the first domain-enhanced semantic vector feature and the second cosine similarity value between the vulnerability attributes with the attribute type of vulnerability description, and determining the sub-similarity value under the vulnerability description by combining the first domain-enhanced semantic vector feature and the second cosine similarity value; Calculating the action and object semantic role annotation result and the third cosine similarity value between the vulnerability attributes with the attribute type of repair suggestion, and determining the sub-similarity value under the repair suggestion by combining the annotation result and the third cosine similarity value; Calculating the classification tree path distance or the second domain-enhanced semantic vector feature between the vulnerability attributes with the attribute type of vulnerability type, and determining the sub-similarity value under the vulnerability type by combining the classification tree path distance or the second domain-enhanced semantic vector feature; Calculating the TF-IDF vector feature and the fourth cosine similarity value between the vulnerability attributes with the attribute type of exploitation method, and determining the sub-similarity value under the exploitation method by combining the TF-IDF vector feature and the fourth cosine similarity value; Calculating the second Jaccard similarity value of the asset type between the vulnerability attributes with the attribute type of impact scope, and determining the sub-similarity value under the impact scope based on the second Jaccard similarity value.

5. The fusion processing method for multi-source vulnerability data according to claim 1, characterized in that After determining whether there is a first type of identifier in the preset mapping knowledge base that has a mapping relationship with the first knowledge base identifier, the multi-source vulnerability data fusion processing method further includes: When there is no first type of identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, using the second type of identifier that matches the first knowledge base identifier as the target knowledge base identifier, and allocating the first vulnerability data under the target knowledge base identifier to obtain the corresponding vulnerability fusion result.

6. The fusion processing method for multi-source vulnerability data according to claim 5, characterized in that The multiple second type of identifiers include multiple second type of identifiers under the first vulnerability knowledge base and multiple second type of identifiers under the second vulnerability knowledge base at the same level as the first vulnerability knowledge base; After it is determined that there is no first type of identifier in the mapping knowledge base that has a mapping relationship with the first knowledge base identifier, the multi-source vulnerability data fusion processing method further includes: When the first knowledge base identifier matches one of the second type identifiers under the first vulnerability knowledge base, determine whether there is a second type identifier under the second vulnerability knowledge base that has a mapping relationship with the first knowledge base identifier in the mapping knowledge base; When there is a second type identifier under the second vulnerability knowledge base that has a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the second type identifier under the second vulnerability knowledge base with the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data to under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

7. The fusion processing method for multi-source vulnerability data according to claim 1, characterized in that After the first knowledge base identifier is matched in order from high to low according to the level order of the pre-arranged multi-level knowledge base identifiers, the multi-source vulnerability data fusion processing method further includes: When the first knowledge base identifier matches one of the first type identifiers, use the first type identifier that matches the first knowledge base identifier as the target knowledge base identifier, and allocate the first vulnerability data to under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

8. The fusion processing method for multi-source vulnerability data according to claim 1, characterized in that The step of allocating the first vulnerability data to under the target knowledge base identifier to obtain a corresponding vulnerability fusion result includes: Determine whether there is a vulnerability storage record of the target knowledge base identifier in the preset vulnerability fusion table; When there is a vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table, update the vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table based on the first vulnerability data to obtain a corresponding vulnerability fusion result; When there is no vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table, obtain the vulnerability description information of the first vulnerability data from the vulnerability knowledge base corresponding to the target knowledge base identifier, and create a vulnerability storage record of the target knowledge base identifier in the vulnerability fusion table based on the vulnerability description information to obtain a corresponding vulnerability fusion result.

9. The fusion processing method for multi-source vulnerability data according to claim 1, characterized in that The multi-source vulnerability data fusion processing method further includes: Determine the vulnerability similarity between the vulnerability data under any one of the first type identifiers and the vulnerability data under any one of the second type identifiers to obtain a corresponding similarity value; When any one of the similarity values is greater than or equal to a preset second similarity threshold, establish a mapping relationship between the first type identifier and the second type identifier corresponding to the determined similar vulnerability data, and save it to the mapping knowledge base.

10. A fusion processing device for multi-source vulnerability data, characterized in that, including: A multi-source data acquisition module, configured to respectively acquire multiple vulnerability data discovered by multiple different channels. Among them, when the vulnerability data carries the knowledge base identifier of the corresponding vulnerability knowledge base, the vulnerability data with the knowledge base identifier is scanned by the corresponding channel based on any one of the multiple vulnerability knowledge bases; An identification matching module, configured to perform fusion processing on each of the vulnerability data in sequence, and for the first vulnerability data currently undergoing fusion processing, if the first vulnerability data carries a corresponding first knowledge base identifier, match the first knowledge base identifier in descending order according to the level order of the pre-arranged multi-level knowledge base identifiers, where the multi-level knowledge base identifiers include a plurality of first type identifiers and a plurality of second type identifiers, and the level of the first type identifiers is higher than that of the second type identifiers; A mapping determination module, configured to determine whether there is a first type identifier having a mapping relationship with the first knowledge base identifier in a preset mapping knowledge base when the first knowledge base identifier does not match any of the first type identifiers but matches one of the second type identifiers; A result determination module, configured to, when there is a first type identifier having a mapping relationship with the first knowledge base identifier in the mapping knowledge base, use the first type identifier having the mapping relationship as the target knowledge base identifier, and allocate the first vulnerability data under the target knowledge base identifier to obtain a corresponding vulnerability fusion result.

11. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method for fusing multi-source vulnerability data according to any one of claims 1 to 10 is implemented.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method for fusing multi-source vulnerability data according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Vulnerability knowledge graph processing method and device, equipment and medium

    CN115827895A

  • Multi-source fusion-based vulnerability knowledge graph construction method

    CN118332492A

  • Vulnerability data analysis method and apparatus, electronic device and storage medium

    WO2024131496A1