Software version inconsistent information detection method for multi-source POC report

By preprocessing and model training of POC reports, identifying and comparing software version information in multiple source POC reports, the problem of inconsistent software version detection in the existing technology is solved, and more reliable vulnerability repair and security protection is achieved.

CN119989364APending Publication Date: 2025-05-13NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073774.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect inconsistent information of affected software versions in multi-source POC reports, resulting in incomplete vulnerability repair and security risks for user choice.

Method used

By preprocessing the POC report, the named entity recognition and relationship extraction model is trained, the name and version information of the affected software are identified, and the inconsistent information of software versions is detected by comparing reports from multiple POC data sources.

Benefits of technology

It realizes automated detection of inconsistent software version information in multi-source POC reports, improves the comprehensiveness of vulnerability repair and reporting quality, and reduces the possibility of users choosing potentially risky versions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989364A_ABST
    Figure CN119989364A_ABST
Patent Text Reader

Abstract

The invention relates to a software version inconsistent information detection method for a multi-source POC report. For a plurality of independent POC data sources, POC reports are preprocessed, and a denoised POC report data set is generated. On the basis, the name, the version and the incidence relation of the affected software are extracted from the POC report by utilizing a named entity identification and relation extraction model, so that the data of different POC reports under the same CVE ID is compared, and an inconsistent result about the version information of the affected software among the data sources is obtained. The method aims to solve the problems of long time consumption and large resource consumption of manual identification of inconsistent information, and can quickly and efficiently detect the inconsistency between affected software versions, thereby improving the comprehensiveness of vulnerability repair, reducing potential security risks, enhancing the quality and credibility of POC reports, and improving the efficiency of vulnerability repair. And more reliable data support is provided for software security protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and is particularly applicable to the field of automated analysis in vulnerability management and assessment. Its purpose is to automatically detect and identify the inconsistencies in the affected software version information in multi-source POC reports, to help software developers improve the comprehensiveness of vulnerability repairs, and to help POC report authors improve the quality of reports, thereby providing more reliable data support for software security protection. Background Art

[0002] In the process of vulnerability discovery and repair, the POC (Proof of Concept) report is a very critical reference. The POC report provides a description or an attack example to illustrate the relevant information of the vulnerability. It can provide researchers with the basis for the existence of a specific security vulnerability, the impact prediction, and the display of potential harm. However, there are currently multiple independent POC data sources, each of which maintains a POC report on CVE. Therefore, for the same vulnerability, there may be different POC reports submitted by multiple authors. Although these reports are for the same CVE ID, the affected software version information provided in them may be different. This leads to the problem of inconsistent affected software versions, which has an impact on multiple links such as software use and repair that cannot be ignored. For example, for software users, the inconsistency problem will mislead users to choose versions with hidden dangers, thereby increasing security risks. For vulnerability repair personnel, the inconsistency problem will mislead them to invest their energy in software versions that do not actually have vulnerabilities, while ignoring versions that actually have vulnerabilities and need to be urgently repaired. Therefore, it is very necessary to detect the inconsistent software version information between POC reports.

[0003] Detecting inconsistent software version information in multiple POC data sources is a major challenge. First, in order to verify a specific software vulnerability, the verifier also needs to install various specific versions of dependent software required for the verification environment according to the guidance of the POC report. This means that in addition to the affected software, the POC report also contains the name and version information of the dependent software, which requires accurate identification of the affected software name and software version from the POC report; secondly, since a single POC report may contain the names of multiple affected software and their corresponding version information, it is necessary to accurately identify the version information matching each affected software from the POC report; finally, since the POC report lacks a unified format, the software name and version information in the report may be expressed in various forms, which requires identification of semantically consistent software names and software versions.

[0004] In this regard, the present invention proposes a method for detecting software version inconsistency information for multi-source POC reports. Based on the preprocessing of the POC reports, the method selects a certain amount of POC reports for annotation, and uses them to train two models to respectively realize named entity recognition and relationship extraction, and obtains the matching results between POC data sources based on this information. First, all POC reports of the POC data source are preprocessed to remove garbled characters, continuous special characters and hyperlinks; secondly, the named entity recognition model is trained with the annotated data to automatically extract the affected software names, affected software versions and CVE IDs as inputs for the subsequent relationship extraction model; thirdly, the relationship extraction model is trained with the annotated data to automatically identify the relationship between the affected software names and software versions; finally, by comparing the POC reports in multiple data sources under the same CVE ID, the software version inconsistency information contained in different matching types of the affected software version information is detected. Summary of the invention

[0005] The present invention provides a software version inconsistency information detection method for multi-source POC reports to effectively evaluate the inconsistency between the affected software version information between POC data sources, thereby expanding the scope of vulnerability repair and reducing the possibility of users selecting software versions with potential risks.

[0006] To achieve the above goals, this method first obtains all POC reports map_all_reports from the POC data source, and generates the preprocessed map_preprocessed by noise removal; secondly, randomly select some POC reports true_dataset from map_preprocessed as annotation objects, obtain annotated_data through annotation, and use it to train the named entity recognition model. After training, the trained named entity recognition model is used to identify the affected software names and affected software versions in the POC report; thirdly, the labeled dataset annotated_data is used to train the relationship extraction model. After training, the trained relationship extraction model is used to identify the relationship between the affected software names and software versions in the POC report; finally, the version_range_keyword_list list is constructed, the software version text description is converted into a set, and finally the detection results of the inconsistent information about the software version are obtained.

[0007] Specifically, the method includes the following steps.

[0008] 1) POC report preprocessing: We first obtain all POC reports from multiple POC data sources and map them in dictionary form map_all_reports <poc_source_name,poc_source_reports_list=<poc1,poc2,...poc n >>Organization, where poc_source_name and poc_source_reports_list correspond to the name of each data source and the corresponding POC report list. In poc_source_reports_list, poc i Refers to the i-th POC report from the poc_source_name data source, where i∈[1,n]; Next, we use regular expressions to remove consecutive special characters, garbled characters, and hyperlinks in each POC report, and finally obtain the preprocessed map_preprocessed <poc_source_name,poc_source_processed_reports_list=<poc_processed1,poc_processed2,...poc_processed n >>, where poc_source_name and poc_source_processed_reports_list refer to the name of the data source and the corresponding pre-processed POC report list, respectively. i Refers to the i-th poc report after preprocessing of the poc_source_name data source result, where i∈[1,n].

[0009] 2) POC report entity recognition: First, randomly select some POC reports from map_preprocessed to generate the dataset true_dataset for annotation, set the annotation target to affected_softname, affected_version entity types and impact, no relationship types, and export the results as annotated_data after the annotation is completed; then, use the BIO tagging method to convert the annotated_data into data_word, divide it into training set train, validation set valid and test set test according to a certain ratio, and use the train dataset to train the named entity recognition model of Embedding+BiLSTM+CRF: the input data is converted into Glove word embedding vector through the Embedding layer, sent to BiLSTM to capture sequence dependency, and then optimized by CRF to predict the label sequence to ensure global optimization, so as to obtain the trained entity_recognition_model; finally, use CVE ID to group map_preprocessed, and use entity_recognition_model to identify the entities in each POC report, and output map_ner <cve_id,list_poc_report_ner=<poc1_ner,poc2_ner,...poc n _ner>>, where poc i _ner refers to the named entity results for cve_id from the i-th data source, which contains the entities and entity types we want (affected software names and affected software versions), where i∈[1,n].

[0010] 3) Entity relationship recognition: We first import the annotated dataset annotated_data, filter out the POC reports annotated with impact and no relationships, extract a pair of entities, the relationship between entities, and the sentences containing the two entities from the annotation results of the POC report and the POC report, and save them to data_re; then, randomly divide data_re into training set train_re, validation set valid_re, and test set test_re according to a certain ratio, and use train_re to train the relationship extraction model that integrates BERT, convolutional layer, and fully connected layer: BERT processes sentence pairs, extracts deep features, and generates contextual representations; the convolutional layer further captures local features in these contextual representations; the fully connected layer performs classification. Thus, the trained relationship extraction model relation_extraction_model is obtained; then, traverse map_ner, and use relation_extraction_model to perform relationship recognition on each key list_poc_report_ner; finally, output map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...poc n _re>>, where poc i _ner refers to the relationship extraction result for cve_id from the i-th data source, where i∈[1,n].

[0011] Table 1 Mapping table of software version text to mathematical range

[0012]

[0013] 4) Software version inconsistency detection: Given map_re, CPE dictionary and version_range_keyword_list for converting software version text description into mathematical range, we first collect text expressions representing version ranges and their corresponding mathematical expressions from POC reports and store them in version_range_keyword_list = [(desc1, alg1), (desc2, alg2) ... (desc n ,alg n )], where desc i and alg iRespectively represent a pair of text descriptions of software versions and corresponding mathematical symbols, where i∈[1,n], as shown in Table 1; at the same time, obtain the CPE dictionary from the NVD official website, and store the software name and the corresponding software version information in the cpe_dic dictionary; then, we set four matching types, namely, full match (two software version sets are completely consistent), loose match (two software version sets are true subsets), full mismatch (two software version sets do not have any identical parts) and cross match (two software version sets have the same content and different content); secondly, traverse each key-value pair in map_re, and convert all software versions in the value into a mathematical range according to the version_range_keyword_list rule. Then, according to the software name in the value, obtain the corresponding software version list from the cpe_dic dictionary, so as to convert the mathematical range into a specific software version set; finally, detect the matching type between multiple POC reports corresponding to each cve_id in map_re, and output the final result, i.e. map_result <cve_id,matrix_match_type n*n >, where matrix_match_type i,j Indicates poc i Data source and POC j Four matching type results of data sources, where i,j∈[1,n], n represents the number of POC data sources. At the same time, the output also includes four_type_result=<matrix_match_rate1,matrix_match_rate2,matrix_match_rate3,matrix_match_rate4> Here, matrix_match_rate1 to matrix_match_rate4 respectively represent the detection results between any two data sources in the following four types, where matrix_match_rate1 represents the detection result about the complete match type, matrix_match_rate2 represents the detection result about the loose match type, matrix_match_rate3 represents the detection result about the complete mismatch type, and matrix_match_rate4 represents the detection result about the cross match type.

[0014] Further, the specific steps of the above step 1) are as follows:

[0015] Step 1)-1: Starting state;

[0016] Step 1)-2: Obtain all original POC reports from each POC data source;

[0017] Step 1)-3: Organize the obtained POC reports in the dictionary map_all_reports, where the name of each data source poc_source_name is used as the key, and the corresponding POC report list poc_source_reports_list = <poc1,poc2,...,poc n >As a value, it indicates all POC reports contained in the data source;

[0018] Step 1)-4: Initialize the regular expression pattern to remove consecutive special characters, garbled characters and hyperlinks in the report;

[0019] Step 1)-5: Clean and preprocess the POC report list of each data source in the dictionary map_all_reports in turn:

[0020] Step 1)-6: For each value of the dictionary map_all_reports, for each POC report poc in poc_source_reports_list i , extract its text content text and prepare for cleaning;

[0021] Step 1)-7: Check whether text contains consecutive special characters, such as "!!!, ###", etc.; if so, use pattern to remove these special characters; if not, execute step 1)-8;

[0022] Step 1)-8: If there are no consecutive special characters in the text, check whether there are garbled characters (such as non-recognizable characters) or hyperlinks (such as links in the format of "http: / / " or "https: / / "); if they exist, use pattern to remove these garbled characters and hyperlinks. If not, execute step 1)-9;

[0023] Step 1)-9: Replace the text after removing special characters, garbled characters and hyperlinks back to the corresponding POC report poc_processed i middle;

[0024] Step 1)-10: Store the preprocessed dictionary result into map_preprocessed;

[0025] Step 1)-11: After the preprocessing process is completed, output map_preprocessed to ensure that all POC reports have been cleaned up;

[0026] Step 1)-12: End state.

[0027] Further, the specific steps of the above step 2) are as follows:

[0028] Step 2)-1: Starting state;

[0029] Step 2)-2: Randomly extract some preprocessed POC reports from map_preprocessed to generate the annotated dataset true_dataset;

[0030] Step 2)-3: Set the recognition target, including entity types affected_softname, affected_version and relationship types impact and no, and label them;

[0031] Step 2)-4: Save the annotated data as annotated_data for subsequent model training;

[0032] Step 2)-5: Convert annotated_data to data_word using BIO notation;

[0033] Step 2)-6: Divide the data according to a certain ratio, and divide data_word into training set train, validation set valid and test set test;

[0034] Step 2)-7: Use the training set train to train the Embedding-BiLSTM-CRF model entity_recognition_model;

[0035] Step 2)-8: The input data is converted into GloVe word embedding vectors through the Embedding layer, and then sent to BiLSTM to capture sequence dependencies. CRF then optimizes label sequence prediction to ensure global optimization.

[0036] Step 2)-9: Group the POC reports in map_preprocessed by CVE ID, and use the trained entity_recognition_model to identify the entities in each POC report one by one;

[0037] Step 2)-10: Generate map_ner <cve_id,list_poc_report_ner=<poc1_ner,poc2_ner,...,poc n _ner>>, where each poc i _ner contains the target entity and type recognition results;

[0038] Step 2)-11: Output map_ner, and the named entity recognition process of all POC reports is completed;

[0039] Step 2)-12: End state.

[0040] Further, the specific steps of the above step 3) are as follows:

[0041] Step 3)-1: Starting state;

[0042] Step 3)-2: Import the annotated dataset annotated_data;

[0043] Step 3)-3: Filter the POC report annotation results containing impact and no relations from annotated_data to extract entity and relationship information;

[0044] Step 3)-4: For each qualified POC report, extract the entity pairs containing the target relationship and the sentences they are in, and save this information as data_re;

[0045] Step 3)-5: Randomly divide data_re into training set, validation set and test set in proportion to ensure the diversity of model training and the reliability of testing;

[0046] Step 3)-6: Train train_re using the relation extraction model consisting of BERT, convolutional layer and fully connected layer;

[0047] Step 3)-7: BERT processes the sentence pairs, extracts deep features, and generates contextual representations; the convolutional layer further captures local features in these contextual representations; the fully connected layer performs classification;

[0048] Step 3)-8: Get the trained relation extraction model relation_extraction_model;

[0049] Step 3)-9: For each POC report in map_ner, if the report contains both the affected software name and the affected software version, use relation_extraction_model to identify the relationship between the entity pairs;

[0050] Step 3)-10: Save the relationship recognition results as map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...,poc n _re>>, where each poc i _re contains the recognition results of the target entity relationship;

[0051] Step 3)-11: The relationship identification process of all POC reports is completed, and the generated map_re is saved as a result file;

[0052] Step 3)-12: End state.

[0053] Further, the specific steps of the above step 4) are as follows:

[0054] Step 4)-1: Starting state;

[0055] Step 4)-2: Get the recognition results of each POC report from map_re, and initialize the software version range conversion list version_range_keyword_list;

[0056] Step 4)-3: According to the software version description in the POC report, expand the version_range_keyword_list, convert each version description into the corresponding mathematical expression, and store the result in the form of (desc, alg) into the version_range_keyword_list list;

[0057] Step 4)-4: Get CPE data from the NVD official website, obtain the software name and corresponding software version list, and store them in the cpe_dic dictionary;

[0058] Step 4)-5: Define four matching types to detect version inconsistency: exact match, loose match, complete mismatch, and cross match;

[0059] Step 4)-6: For each cve_id in map_re, extract all POC reports corresponding to the vulnerability and obtain the software version information involved;

[0060] Step 4)-7: Use version_range_keyword_list to convert the software version description in the POC report into a mathematical range;

[0061] Step 4)-8: Use the cpe_dic dictionary to map the converted mathematical range to a specific version set and generate a version list for inconsistency matching;

[0062] Step 4)-9: For each POC report under cve_id, compare the matching types of each report software version set and record the result to four_type_result=<matrix_match_rate1,matrix_match_rate2,matrix_match_rate3,matrix_match_rate4> , where matrix_match_rate1 to matrix_match_rate4 respectively contain the detection results of four types of complete match, loose match, complete mismatch and cross match between any pair of data sources;

[0063] Step 4)-10: Organize the four types of matrices into map_result by cve_id <cve_id,matrix_match_type n*n >Dictionary, which is used to record the four types of inconsistency detection in the POC report version set corresponding to each CVE ID.

[0064] Step 4)-11: Output the final results of the inconsistency analysis, map_result and four_type_result;

[0065] Step 4)-12: End state. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 The present invention is a flowchart of a method for detecting software version inconsistency information in multi-source POC reports in an implementation of the present invention.

[0067] Figure 2 for Figure 1 Flowchart of POC report preprocessing generation.

[0068] Figure 3 for Figure 1 Flowchart of entity recognition generation in the POC report.

[0069] Figure 4 for Figure 1 Flowchart of entity relationship recognition in .

[0070] Figure 5 for Figure 1 Flowchart generated by software version inconsistency detection. DETAILED DESCRIPTION

[0071] In order to better understand the technical content of the present invention, a specific implementation is given and described as follows with reference to the accompanying drawings.

[0072] Figure 1The present invention is a flowchart of a method for detecting software version inconsistency information in multi-source POC reports in an implementation of the present invention.

[0073] A method for detecting software version inconsistency information based on multi-source POC reports, characterized by comprising the following steps.

[0074] S1 POC report preprocessing, this step extracts all POC reports and organizes them into a dictionary map_all_reports, where the name of each data source poc_source_name is used as the key and the corresponding POC report poc_source_reports_list is in list form as the value; then, regular expressions are used to remove consecutive special characters, garbled characters, and hyperlinks in the report to generate the preprocessed map_preprocessed <poc_source_name,poc_source_processed_reports_list=<poc_processed1,poc_processed2,...poc_processed n >>Dictionary.

[0075] S2 POC report entity recognition. This step randomly selects reports from map_preprocessed to generate true_dataset, and generates annotated_data after annotation. Then, the entity information is extracted from annotated_data and converted into BIO notation format to generate data_word. Then, data_word is divided into training set, validation set and test set according to a certain ratio, and the entity recognition model is trained with the training set to obtain entity_recognition_model. Finally, the reports in map_preprocessed are traversed, entity_recognition_model is used to identify entity types, and the results are saved to map_ner.

[0076] S3 entity relationship recognition, this step imports annotated_data, filters reports marked with impact or no relationships, and extracts text, a pair of entity location information, a pair of entity content entity1 and entity2, relationship relation, and sentences containing entity1 and entity2, and stores them in data_re. Then, data_re is divided into training set train_re, validation set valid_re, and test set test_re in a certain proportion, and the training set train_re is used to train the relationship extraction model composed of Bert, convolutional layer, and fully connected layer to obtain relation_extraction_model. Finally, traverse map_ner, extract entities and sentences, input them into relation_extraction_model to identify entity relationships, and save the results to map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...poc n _re>>among.

[0077] S4 software version inconsistency detection, this step collects text descriptions and mathematical expressions representing version ranges from POC reports, and stores them in version_range_keyword_list, and obtains CPE dictionary from NVD official website and stores them in cpe_dic dictionary. Then, traverse each key-value pair in map_re, and convert all software versions in the value into mathematical ranges according to version_range_keyword_list rules; again, according to the corresponding software version obtained from cpe_dic dictionary, convert the mathematical range into a specific software version set; finally, detect the matching type between the POC reports corresponding to each cve_id in map_re, and output the final results map_result and four_type_result.

[0078] Figure 2Flowchart generated for POC report preprocessing. All POC reports are extracted and organized in the dictionary map_all_reports, where the name of each data source poc_source_name is used as the key and the corresponding POC report poc_source_reports_list is in list form as the value; then, regular expressions are used to remove consecutive special characters, garbled characters, and hyperlinks in the report to generate the preprocessed map_preprocessed <poc_source_name,poc_source_processed_reports_list=<poc_processed1,poc_processed2,...,poc_processed n >>Dictionary.

[0079] The specific steps are as follows:

[0080] Step 1: Starting state; Step 2: Get all original POC reports from each POC data source; Step 3: Organize the obtained POC reports in the dictionary map_all_reports, where the name of each data source poc_source_name is used as the key, and the corresponding POC report list poc_source_reports_list = <poc1,poc2,...,poc n > as a value, indicating all POC reports contained in the data source; Step 4: Initialize the regular expression pattern to remove consecutive special characters, garbled characters and hyperlinks in the report; Step 5: For the POC report list of each data source in the dictionary map_all_reports, clean and preprocess them in turn: Step 6: For each value of the dictionary map_all_reports, poc_source_reports_list, each POC report poc i , extract its text content text, prepare to clean; Step 7: Check whether the text contains continuous special characters, such as "!!!, ###", etc.; if so, use pattern to remove these special characters, if not, execute steps 1)-8; Step 8: If there are no continuous special characters in the text, check whether there are garbled characters (such as non-recognizable characters) or hyperlinks (such as links in the format of "http: / / " or "https: / / "); if so, use pattern to remove these garbled characters and hyperlinks, if not, execute step 9; Step 9: Replace the text after removing special characters, garbled characters and hyperlinks with the corresponding POC report poc_processed iStep 10: Store the preprocessed dictionary result in map_preprocessed; Step 11: After the preprocessing process is completed, output map_preprocessed to ensure that all POC reports have been cleaned up; Step 12: End state.

[0081] Figure 3 Flowchart generated for entity recognition of POC reports. Randomly select reports from map_preprocessed to generate true_dataset, and generate annotated_data after annotation; then, extract entity information from annotated_data and convert it into BIO notation format to generate data_word; then, divide data_word into training set, validation set and test set according to a certain ratio, use the training set to train the entity recognition model, and get entity_recognition_model; finally, traverse the reports in map_preprocessed, use entity_recognition_model to identify entity types, and save the results to map_ner. The specific steps are as follows:

[0082] Step 1: Starting state; Step 2: Randomly extract some preprocessed POC reports from map_preprocessed to generate the annotated data set true_dataset; Step 3: Set the recognition target, including entity types affected_softname, affected_version and relationship types impact and no, for annotation; Step 4: Save the annotated data as annotated_data for subsequent model training; Step 5: Convert annotated_data to data_word using BIO notation; Step 6: Divide data according to a certain ratio, and divide data_word into training set train, validation set valid and test set test; Step 7: Use the training set train to train the Embedding-BiLSTM-CRF model entity_recognition_model; Step 8: The input data is converted into Glove word embedding vectors through the Embedding layer, sent to BiLSTM to capture sequence dependencies, and then optimized by CRF to predict the label sequence to ensure global optimization; Step 9: The POC report in map_preprocessed is converted according to CVE ID grouping, and use the trained entity_recognition_model to identify the entities in each POC report one by one; Step 10: Generate map_ner <cve_id,list_poc_report_ner=<poc1_ner,poc2_ner,...,pocn _ner>>, where each poc i _ner contains the target entity and type recognition results; Step 11: Output map_ner, and the named entity recognition process of all POC reports is completed; Step 12: End state.

[0083] Figure 4 Flowchart generated for entity relationship recognition. Import annotated_data, filter reports marked with impact or no relationships, and extract text, a pair of entity location information, a pair of entity content entity1 and entity2, relationship relation, and sentence sentence containing entity1 and entity2, and store them in data_re. Then, divide data_re into training set train_re, validation set valid_re, and test set test_re in a certain proportion, and use the training set train_re to train the relationship extraction model composed of Bert, convolutional layer, and fully connected layer to obtain relation_extraction_model. Finally, traverse map_ner, extract entities and sentences, input them into relation_extraction_model to identify entity relationships, and save the results to map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...,poc n _re>>. The specific steps are as follows:

[0084] Step 1: Starting state; Step 2: Import the annotated dataset annotated_data; Step 3: Filter the POC report annotation results containing impact and no relations from annotated_data to extract entity and relationship information; Step 4: For each qualified POC report, extract the entity pairs containing the target relationship and the sentences they are in, and save this information as data_re; Step 5: Randomly divide data_re into training set, validation set and test set in proportion to ensure the diversity of model training and the reliability of test; Step 6: Use the relationship composed of BERT, convolutional layer and fully connected layer Extraction model trains train_re; Step 7: BERT processes sentence pairs, extracts deep features, and generates contextual representations; the convolutional layer further captures local features in these contextual representations; the fully connected layer performs classification; Step 8: Get the trained relation extraction model relation_extraction_model; Step 9: For each POC report in map_ner, if the report contains both the affected software name and the affected software version, use relation_extraction_model to identify the relationship between the entity pairs; Step 10: Save the relationship identification result as map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...,poc n _re>>, where each poc i _re contains the recognition results of the target entity relationship; Step 11: The relationship recognition process of all POC reports is completed, and the generated map_re is saved as a result file; Step 12: End state.

[0085] Figure 5 Flowchart generated for software version inconsistency detection. Collect text descriptions and mathematical expressions representing the version range from the POC report and store them in version_range_keyword_list, obtain the CPE dictionary from the NVD official website, and store it in the cpe_dic dictionary. Then, traverse each key-value pair in map_re, and convert all software versions in the value into a mathematical range according to the version_range_keyword_list rule; again, according to the corresponding software version obtained from the cpe_dic dictionary, convert the mathematical range into a specific set of software versions; finally, detect the matching type between the POC reports corresponding to each cve_id in map_re, and output the final results map_result and four_type_result. The specific steps are as follows:

[0086] Step 1: Starting state; Step 2: Get the recognition results of each POC report from map_re, and initialize the software version range conversion list version_range_keyword_list; Step 3: According to the software version description in the POC report, expand version_range_keyword_list, convert each version description into the corresponding mathematical expression, and store the result in the form of (desc, alg) in the version_range_keyword_list list; Step 4: Get CPE data from the NVD official website, get the software name and the corresponding software version list, and store it in the cpe_dic dictionary; Step 5: Define four types Matching type to detect version inconsistency: full match, loose match, full mismatch and cross match; Step 6: For each cve_id in map_re, extract all POC reports corresponding to the vulnerability and obtain the software version information involved; Step 7: Use version_range_keyword_list to convert the software version description in the POC report into a mathematical range; Step 8: Use the cpe_dic dictionary to map the converted mathematical range to a specific version set and generate a version list for inconsistency matching; Step 9: For each POC report under cve_id, compare the matching type of each report software version set and record the result to four_type_result=<matrix_match_rate1,matrix_match_rate2,matrix_match_rate3,matrix_match_rate4> , where matrix_match_rate1 to matrix_match_rate4 contain the detection results of four types of complete match, loose match, complete mismatch and cross match between any pair of data sources in turn; Step 10: Organize the four types of matrices into map_result according to cve_id <cve_id,matrix_match_type n*n >Dictionary, which is used to record the inconsistency detection of these four types in the POC report version set corresponding to each CVE ID. Step 11: Output the final results of the inconsistency analysis, map_result and four_type_result; Step 12: End status.

[0087] In summary, the present invention solves the problem of detecting inconsistent POC report information between multiple POC data sources. By acquiring and preprocessing POC reports, annotating and training named entity recognition and relationship extraction models, and converting software version text descriptions, the present invention finally obtains the detection results of inconsistent software version information.

Claims

1. A software version inconsistency information detection method for multi-source POC reports, characterized in that: This method first obtains all POC reports map_all_reports from the POC data source, and generates the preprocessed map_preprocessed by noise removal; secondly, randomly selects part of the POC report true_dataset from map_preprocessed as the annotation object, obtains annotated_data through annotation, and uses it to train the named entity recognition model. After training, the trained named entity recognition model is used to identify the affected software names and affected software versions in the POC report; thirdly, the annotated dataset annotated_data is used to train the relationship extraction model. After training, the trained relationship extraction model is used to identify the relationship between the affected software names and software versions in the POC report; finally, the version_range_keyword_list list is constructed, the software version text description is converted into a set, and finally the detection result of the software version inconsistency information is obtained; the method includes the following steps: 1) POC report preprocessing: We first obtain all POC reports from multiple POC data sources and map these POC reports in dictionary form map_all_reports <poc_source_name,poc_source_reports_list=<poc1,poc2,...poc n >>Organization, where poc_source_name and poc_source_reports_list correspond to the name of each data source and the corresponding POC report list. In poc_source_reports_list, poc i Refers to the i-th POC report from the poc_source_name data source, where i∈[1,n]; Next, we use regular expressions to remove consecutive special characters, garbled characters, and hyperlinks in each POC report to obtain the preprocessed map_preprocessed <poc_source_name,poc_source_processed_reports_list=<poc_processed1,poc_processed2,...poc_processed n >>, where poc_source_name and poc_source_processed_reports_list refer to the name of the data source and the corresponding pre-processed POC report list, respectively. i Refers to the i-th poc report after preprocessing of the poc_source_name data source result, where i∈[1,n]; 2) POC report entity recognition: First, randomly select some POC reports from map_preprocessed to generate the dataset true_dataset for annotation, set the annotation target to affected_softname, affected_version entity types and impact, no relationship types, and export the results as annotated_data after the annotation is completed; then, use the BIO tagging method to convert the annotated_data into data_word, divide it into training set train, validation set valid and test set test according to a certain ratio, and use the train dataset to train the named entity recognition model of Embedding+BiLSTM+CRF: the input data is converted into Glove word embedding vector through the Embedding layer, sent to BiLSTM to capture sequence dependency, and then optimized by CRF to predict the label sequence to ensure global optimization, so as to obtain the trained entity_recognition_model; finally, use CVE ID to group map_preprocessed, and use entity_recognition_model to identify the entities in each POC report, and output map_ner <cve_id,list_poc_report_ner=<poc1_ner,poc2_ner,...poc n _ner>>, where poc i _ner refers to the named entity results for cve_id from the i-th data source, which contains the entities and entity types we want (affected software names and affected software versions), where i∈[1,n]; 3) Entity relationship recognition: We first import the annotated dataset annotated_data, filter out the POC reports annotated with impact and no relationships, extract a pair of entities, the relationship between entities, and the sentences containing the two entities from the annotation results of the POC report and the POC report, and save them to data_re; then, randomly divide data_re into training set train_re, validation set valid_re and test set test_re according to a certain ratio, and use train_re to train the relationship extraction model that integrates BERT, convolutional layer and fully connected layer: BERT processes sentence pairs, extracts deep features, and generates contextual representations; the convolutional layer further captures local features in these contextual representations; the fully connected layer performs classification; thus, a trained relationship extraction model relation_extraction_model is obtained; then, map_ner is traversed, and relation_extraction_model is used to perform relationship recognition on each key list_poc_report_ner; finally, map_re is output <cve_id,list_poc_report_re=<poc1_re,poc2_re,...poc n _re>>, where poc i _ner refers to the relationship extraction result for cve_id from the i-th data source, where i∈[1,n]; 4) Software version inconsistency detection: Given map_re, CPE dictionary and version_range_keyword_list for converting software version text description into mathematical range, we first collect text expressions representing version ranges and their corresponding mathematical expressions from POC reports and store them in version_range_keyword_list = [(desc1, alg1), (desc2, alg2) ... (desc n ,alg n )], where desc i and alg i Respectively represent a pair of text descriptions of software versions and corresponding mathematical symbols, where i∈[1,n], as shown in Table 1; at the same time, obtain the CPE dictionary from the NVD official website, and store the software name and the corresponding software version information in the cpe_dic dictionary; then, we set four matching types, namely, full match (two software version sets are completely consistent), loose match (two software version sets are true subsets), full mismatch (two software version sets do not have any identical parts) and cross match (two software version sets have the same content and different content); secondly, traverse each key-value pair in map_re, and convert all software versions in the value into a mathematical range according to the version_range_keyword_list rule. Then, according to the software name in the value, obtain the corresponding software version list from the cpe_dic dictionary, so as to convert the mathematical range into a specific software version set; finally, detect the matching type between multiple POC reports corresponding to each cve_id in map_re, and output the final result, i.e. map_result <cve_id,matrix_match_type n*n >, where matrix_match_type i,j Indicates POC i Data source and POC j Four matching type results of the data source, where i,j∈[1,n], n represents the number of POC data sources; at the same time, the output also includes four_type_result=<matrix_match_rate1,matrix_match_rate2,matrix_match_rate3,matrix_match_rate4> Here, matrix_match_rate1 to matrix_match_rate4 respectively represent the detection results between any two data sources in the following four types, where matrix_match_rate1 represents the detection result about the complete match type, matrix_match_rate2 represents the detection result about the loose match type, matrix_match_rate3 represents the detection result about the complete mismatch type, and matrix_match_rate4 represents the detection result about the cross match type.

2. The method for detecting inconsistency information in a POC report according to claim 1, characterized in that: In step 1), all POC reports are extracted and organized in the dictionary map_all_reports, where the name of each data source poc_source_name is used as the key and the corresponding POC report poc_source_reports_list is in list form as the value; then, regular expressions are used to remove consecutive special characters, garbled characters and hyperlinks in the report to generate the preprocessed map_preprocessed <poc_source_name,poc_source_processed_reports_list=<poc_processed1,poc_processed2,...poc_processed n >>Dictionary.

3. The method for detecting inconsistency information in a POC report according to claim 1, characterized in that: In step 2), a report is randomly selected from map_preprocessed to generate true_dataset, and annotated_data is generated after annotation; then, entity information is extracted from annotated_data and converted into BIO notation format to generate data_word; then, data_word is divided into training set, validation set and test set according to a certain ratio, and the entity recognition model is trained with the training set to obtain entity_recognition_model; finally, the reports in map_preprocessed are traversed, entity_recognition_model is used to identify entity types, and the results are saved to map_ner.

4. The method for detecting inconsistency information in a POC report according to claim 1, characterized in that: In step 3), import annotated_data, filter reports marked with impact or no relations, extract text, a pair of entity location information, a pair of entity contents entity1 and entity2, relation relation, and sentence sentence containing entity1 and entity2, and store them in data_re; then, divide data_re into training set train_re, validation set valid_re, and test set test_re in a certain proportion, and use the training set train_re to train the relation extraction model composed of Bert, convolutional layer, and fully connected layer to obtain relation_extraction_model; finally, traverse map_ner, extract entities and sentences, input them into relation_extraction_model to identify entity relations, and save the results to map_re <cve_id,list_poc_report_re=<poc1_re,poc2_re,...poc n _re>>among.

5. The method for detecting inconsistency information in a POC report according to claim 1, characterized in that: In step 4), collect text descriptions and mathematical expressions representing version ranges from POC reports and store them in version_range_keyword_list, obtain CPE dictionary from NVD official website and store them in cpe_dic dictionary; then, traverse each key-value pair in map_re, and convert all software versions in the value into mathematical ranges according to version_range_keyword_list rules; again, convert the mathematical range into a specific set of software versions according to the corresponding software versions obtained from cpe_dic dictionary; finally, detect the matching type between the POC reports corresponding to each cve_id in map_re, and output the final results map_result and four_type_result.