Software vulnerability detection method and device, equipment and storage medium
By combining port scanning and web service identification tools, time features and multi-source features fusion of Transformer models, the accuracy and efficiency of vulnerability detection in the existing technology are solved, and more accurate software vulnerability positioning and risk assessment are achieved.
Patent Information
- Application Number
- CN202510787383.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-12
AI Technical Summary
现有技术在软件漏洞检测中存在漏洞匹配准确率低、版本范围修正能力不足,无法有效利用时间特征和动态信息,导致漏洞检测效率低下。
The initial software version interval is determined by combining the port scanning tool and the web service identification tool, and the file change time and version release time are used to correct it; multi-source feature fusion and training are used to generate version probability distribution; and the vulnerability test script is tested and adjusted based on the modified version interval, and the probability of the vulnerability exists is finally determined.
It improves the accuracy and efficiency of vulnerability detection, can quickly locate risks, reduce system security risks, and provide accurate vulnerability probability and risk assessment data support.
Smart Images

Figure CN120296750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and particularly to a software vulnerability detection method, device, equipment, and storage medium. Background Art
[0002] In the field of network security, vulnerability scanning and risk assessment are crucial for resisting network attacks, and accurately identifying the software version of the target system is the key to matching vulnerabilities.
[0003] However, there are many defects in the prior art. On the one hand, the fuzz testing and version analysis technology based on large language models focuses on generating test cases to improve the vulnerability detection coverage rate. Its core logic is to "discover the existence of vulnerabilities", resulting in fuzzy version positioning, only being able to associate to a version range, with low vulnerability matching accuracy and easy generation of false positives and false negatives. On the other hand, this technology relies on a single data source, lacks multi-source feature fusion, and cannot utilize dynamic information such as time features and POC test (i.e., Proof of Concept, verification test) results, resulting in insufficient version range correction ability, decoupling of POC verification and version range, inability to effectively narrow the version range, and failure to incorporate time features into the analysis, resulting in redundant version ranges and low vulnerability matching efficiency.
[0004] Therefore, how to improve the accuracy of vulnerability detection is an urgent technical problem to be solved currently. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a software vulnerability detection method, device, equipment, and storage medium, which can improve the accuracy of vulnerability detection. The specific solutions are as follows:
[0006] In the first aspect, the present application provides a software vulnerability detection method, including:
[0007] Using a port scanning tool and a Web service identification tool to determine the initial software version range of the target software from the target system, and correcting the initial software version range based on the file change time of the target file and the version release time of the target software to obtain a corrected version range; wherein, the target file is included in the target software;
[0008] Training a pre-trained model based on the Transformer architecture using a preset vulnerability database, and determining a version probability distribution based on the trained target large language model;
[0009] Determining a target vulnerability test script from the preset vulnerability database based on the corrected version range, and testing the target system based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range;
[0010] Determine the vulnerability existence probability in the target system based on the version probability distribution and the target version interval.
[0011] Optionally, determining the initial software version interval of the target software from the target system by using the port scanning tool and the Web service identification tool includes:
[0012] Use the port scanning tool and the Web service identification tool respectively to determine the network protocol identifier, response header information, and service type of the target software from the target system;
[0013] Determine the first software version interval through the port scanning tool based on the network protocol identifier, the response header information, and the service type;
[0014] Determine the second software version interval through the Web service identification tool based on the network protocol identifier, the response header information, and the service type;
[0015] Determine the intersection interval of the first software version interval and the second software version interval, and determine the intersection interval as the initial software version interval.
[0016] Optionally, correcting the initial software version interval based on the file change time of the target file and the version release time of the target software to obtain the corrected version interval includes:
[0017] Determine the version release time of the target software from the preset vulnerability database and the preset manufacturer information;
[0018] Determine the target static file from the target software, and use the preset Hypertext Transfer Protocol request to determine the file change time of the target static file closest to the current time;
[0019] Determine the version timeline based on the version release time and the file change time, and use the version timeline to correct the initial software version interval to obtain the corrected version interval.
[0020] Optionally, training a pre-trained model based on the Transformer architecture using the preset vulnerability database, and determining the version probability distribution based on the trained target large language model includes:
[0021] Configure a multi-source feature encoding layer for the input layer of the pre-trained model based on the Transformer architecture, and configure a normalization exponential function for the output layer of the pre-trained model to obtain the configured model;
[0022] Obtain training data from the preset vulnerability database and the preset vendor information, and train the configured model based on the training data to obtain a target large language model;
[0023] Determine the website request return information of the target software and the time information of the target static file obtained from the target system by using the port scanning tool and the Web service identification tool to determine the first input information;
[0024] Determine the version usage data and version priority of the target software as the second input information;
[0025] Input the first input information and the second input information into the target large language model so that the target large language model outputs a version probability distribution.
[0026] Optionally, testing the target system based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range, including:
[0027] Determine an attack payload based on the target vulnerability test script and send the attack payload to the target system to obtain a system response result;
[0028] Match the system response result with preset vulnerability characteristics to obtain a matching result;
[0029] If the matching result indicates that there is a vulnerability in the target system, determine the scope of influence of the vulnerability in the target system;
[0030] Adjust the corrected version range based on the scope of influence of the vulnerability to obtain a target version range.
[0031] Optionally, after determining the probability of the existence of a vulnerability in the target system based on the version probability distribution and the target version range, it further includes:
[0032] Search for a target vulnerability from the preset vulnerability database based on the target version range;
[0033] Determine the version coverage range of the target vulnerability and determine the probability of the existence of a vulnerability in the target system based on the version probability distribution and the version coverage range.
[0034] Optionally, after determining the probability of the existence of a vulnerability in the target system based on the version probability distribution and the target version range, it further includes:
[0035] Score the target vulnerability using a preset common vulnerability scoring system to obtain a scoring result;
[0036] Determine the exploitation status of the target vulnerability; the exploitation status indicates whether there is exploitation code for the target vulnerability;
[0037] Generate a corresponding risk level based on the scoring result and the exploitation status, and perform a corresponding preset risk response operation using the risk level.
[0038] In a second aspect, the present application provides a software vulnerability detection device, including:
[0039] A first interval determination module, configured to determine an initial software version interval of a target software from a target system using a port scanning tool and a Web service identification tool, and correct the initial software version interval based on the file change time of a target file and the version release time of the target software to obtain a corrected version interval; wherein, the target file is included in the target software;
[0040] A probability determination module, configured to train a pre-trained model based on a Transformer architecture using a preset vulnerability database, and determine a version probability distribution based on the trained target large language model;
[0041] A second interval determination module, configured to determine a target vulnerability test script from the preset vulnerability database based on the corrected version interval, and test the target system based on the target vulnerability test script to adjust the corrected version interval using the obtained test result to obtain a target version interval;
[0042] A risk determination module, configured to determine the probability of existence of vulnerabilities in the target system based on the version probability distribution and the target version interval.
[0043] In a third aspect, the present application provides an electronic device, including:
[0044] A memory, configured to store a computer program;
[0045] A processor, configured to execute the computer program to implement the foregoing software vulnerability detection method.
[0046] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the foregoing software vulnerability detection method is implemented.
[0047] In this application, a port scanning tool and a web service identification tool are used to determine the initial software version range of the target software from the target system, and the initial software version range is corrected based on the file change time of the target file and the version release time of the target software to obtain a corrected version range; wherein, the target file is included in the target software; a pre-trained model based on the Transformer architecture is trained using a preset vulnerability database, and a target large language model is used to determine the version probability distribution; a target vulnerability test script is determined from the preset vulnerability database based on the corrected version range, and the target system is tested based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range; the probability of the existence of vulnerabilities in the target system is determined based on the version probability distribution and the target version range. As can be seen from the above, in the process of software vulnerability detection in this application, first, a port scanning tool and a web service identification tool are used to analyze the target system to obtain the initial software version range of the target software. On this basis, the initial software version range is further optimized and corrected according to the file change time of the target file and the version release time of the target software, and then a more accurate corrected version range is obtained. Next, a preset vulnerability database is used to train a pre-trained model based on the Transformer architecture. Through the training process, the model continuously learns and masters the correlation between software versions and various features, and finally forms a target large language model, and the target large language model is used to determine the version probability distribution of the target software. Then, based on the corrected version range, an appropriate target vulnerability test script is selected from the preset vulnerability database. The target system is tested using the target vulnerability test script, and based on the test results, the corrected version range is dynamically adjusted again to obtain a more practical target version range. Finally, the actual probability of the existence of each vulnerability in the target system is determined through the version probability distribution and the target version range, providing certain data support for subsequent vulnerability risk assessment and handling. In this way, this application can improve the efficiency of vulnerability detection, quickly locate risks, and reduce system security risks. Description of the Drawings
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0049] Figure 1 Flowchart of a software vulnerability detection method disclosed in this application;
[0050] Figure 2 A flowchart of a specific software vulnerability detection method disclosed in this application;
[0051] Figure 3 A schematic structural diagram of a software vulnerability detection device disclosed in this application;
[0052] Figure 4 A structural diagram of an electronic device disclosed in this application. Specific implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Currently, there are many defects in the prior art. On the one hand, the fuzz testing and version analysis technology based on large language models focuses on generating test cases to improve the vulnerability detection coverage rate. Its core logic is to "discover the existence of vulnerabilities", resulting in fuzzy version positioning, only being able to associate with a version interval, with low vulnerability matching accuracy and easy to generate false positives and false negatives. On the other hand, this technology relies on a single data source, lacks multi-source feature fusion, and cannot utilize dynamic information such as time features and POC test (i.e., Proof of Concept, verification test) results, resulting in insufficient version range correction ability, decoupling of POC verification and version range, inability to effectively narrow the version interval, and lack of inclusion of time features in the analysis, resulting in redundant version intervals and low vulnerability matching efficiency. For this reason, this application provides a software vulnerability detection method, device, equipment, and storage medium, which can improve the accuracy of vulnerability detection.
[0055] See Figure 1 As shown, the embodiments of the present invention disclose a software vulnerability detection method, including:
[0056] Step S11: Use a port scanning tool and a Web service identification tool to determine the initial software version interval of the target software from the target system, and correct the initial software version interval based on the file change time of the target file and the version release time of the target software to obtain a corrected version interval; wherein, the target file is included in the target software.
[0057] In this embodiment, it can be understood that in the actual scenario of software vulnerability detection, due to the complexity and diversity of the target system, a single detection tool often cannot accurately determine the version of the target software. Therefore, a port scanning tool and a Web service identification tool can be comprehensively used to obtain more comprehensive information about the target software.
[0058] Specifically, when the port scanning tool scans the target system, it will detect the network protocol identifier, response header information, and service type of the open port. For example, when Nmap (a port scanning tool) scans a certain server, Nmap will, based on the received network protocol identifier, which may include protocol fingerprints. At the same time, Nmap will parse the HTTP response header (i.e., the information returned by the server to the client based on the Hypertext Transfer Protocol) information, such as "Server: Apache / 2.4.x", which initially provides the information that the target software is Apache (a Web server software) and the version is in the 2.4 series. In addition, the port scanning tool can also identify the service type, such as Web service, etc. These information are of important reference value for determining the type and version range of the target software.
[0059] Meanwhile, starting from the Web service level, the Web service identification tool analyzes the page elements, script features, etc. of the target system. For example, Wappalyzer (a Web service identification tool) will, based on the information such as library files and script codes referenced in the page, judge the type and version of the target software. For example, by analyzing the version information of a specific library referenced in the page and combining other page element features, Wappalyzer can confirm the version range of the target software. When analyzing the above-mentioned server, Wappalyzer may, based on the features related to Apache in the page, further determine that the version of the Apache software is in the 2.4 series.
[0060] After obtaining the relevant information by using the port scanning tool and the Web service identification tool respectively, next, the initial software version range of the target software is determined. The specific operation is to compare the first software version range obtained by the port scanning tool with the second software version range obtained by the Web service identification tool, and take the intersection of the two. For example, if the first software version range determined by the port scanning tool is v2.4.30 - 2.4.50, and the second software version range determined by the Web service identification tool is v2.4.x, then by taking the intersection, the initial software version range can be obtained as v2.4.30 - 2.4.50. This way of comprehensively analyzing the detection results of multiple tools effectively integrates the detection advantages of different tools, avoids the limitations of a single tool, and significantly narrows the version range.
[0061] After obtaining the initial software version range, in order to further improve the accuracy of version determination, it is necessary to correct it by combining information in the time dimension. This process is mainly achieved by analyzing the version release time of the target software and the file change time of the target file.
[0062] First, collect the release times of each version of the target software from the preset vulnerability database and the official release channels of the manufacturer. Among them, the preset vulnerability database can be the CVE (Common Vulnerabilities and Exposures) database. Taking Apache software as an example, in the CVE database, it can be clearly seen that the 2.4.49 version was released on September 1, 2021, and the 2.4.50 version was released on October 15, 2021. These accurate version release time information constructs the version timeline of the target software.
[0063] At the same time, analyze the static files in the target software, and use the preset Hypertext Transfer Protocol request to determine the file change time of the target static file closest to the current time. Specifically, obtain the Last-Modified (an HTTP response header field used to indicate the last modification time of the resource on the server) timestamp of the file by sending an HTTP request, so as to determine the latest change time of the file. In actual detection, assume that when scanning the target system, it is found that the last modification time of the target file is September 10, 2021. Combining the constructed version timeline, since the change time of this file is after the release of the Apache 2.4.49 version, it can be inferred that the lower limit of the version of the target software is 2.4.49.
[0064] Finally, based on the version timeline determined by the version release time and the file change time, correct the initial software version range, so as to obtain a more accurate corrected version range. For example, if the initial software version range is v2.4.30 - 2.4.50, after the above analysis, the corrected version range can be narrowed down to v2.4.49 - 2.4.50. This process makes full use of time characteristics and dynamically adjusts the version range, making the version range more in line with the actual situation of the target software.
[0065] Step S12: Train the pre-trained model based on the Transformer architecture using the preset vulnerability database, and determine the version probability distribution based on the trained target large language model.
[0066] In this embodiment, first, the structure configuration of the pre-trained model based on the Transformer architecture is optimized. Among them, the pre-trained model can be a BERT model (i.e., Bidirectional Encoder Representations from Transformers, a pre-trained language model). A multi-source feature encoding layer is configured in the input layer of the model. This encoding layer can effectively integrate and encode feature information from different channels with different formats. For example, for network protocol identifiers, response header information, etc. obtained from port scanning tools, and page element features obtained from web service identification tools, the multi-source feature encoding layer can convert them into vector forms that the model can understand, realizing the unified expression of multi-source data. A normalization exponential function (i.e., Softmax) is configured in the output layer. This function can convert the original numerical values output by the model into a probability distribution form, making the output results more interpretable and practical, and facilitating the intuitive judgment of the possibilities of each version. After the above configuration, the configured model is obtained.
[0067] Next, training data is obtained from a preset vulnerability database and preset vendor information. A large number of records related to software versions and vulnerabilities are stored in the preset vulnerability database. These records contain information such as the vulnerability characteristics of the software in different versions, while the preset vendor information provides data such as detailed descriptions of software version releases and version update content. These data are sorted and labeled to form a training data set covering the "version - feature" correspondence relationship. Based on this training data set, the configured model is trained. During the training process, the model continuously adjusts its internal parameters to learn the potential relationship between multi-source features and software versions. After multiple rounds of iterative training, a target large language model that can accurately capture the relationship between features and versions is finally obtained.
[0068] Furthermore, the information for inputting into the target large language model is determined. On the one hand, website request return information about the target software obtained from the target system using port scanning tools and web service identification tools is collected. This information contains various details exposed by the target software during network interaction, such as software identifiers in the response header and software characteristics reflected by page elements. At the same time, the time information of the target static file is obtained. This time information reflects the change situation of the files in the target software and is closely related to the software version. These information are integrated and determined as the first input information. On the other hand, the version usage data and version priority of the target software are determined as the second input information. The version usage data can be obtained from channels such as market research and user statistics, reflecting the popularity of different versions in actual applications; the version priority is determined according to factors such as the software vendor's maintenance strategy for the version and vulnerability repair situation. For example, a version that fixes major security vulnerabilities has a higher priority.
[0069] After obtaining the first input information and the second input information, input the first input information and the second input information into the target large language model. It should be noted that if a certain version fixes major vulnerabilities, then increase its probability weight, and the situation of increasing the probability weight can be regarded as the version priority. The target large language model will comprehensively consider the software feature clues provided by multi-source features, the actual application situation reflected by the version usage volume, and the importance reflected by the version priority. Through internal calculations and processing, it will finally output the version probability distribution. For example, the output result may be {v2.4.49: 65%, v2.4.50: 35%}. This probability distribution clearly shows the likelihood of each version, which helps to perform vulnerability matching and risk assessment work more accurately.
[0070] Step S13: Determine the target vulnerability test script from the preset vulnerability database based on the corrected version range, and test the target system based on the target vulnerability test script, so as to adjust the corrected version range using the obtained test results to obtain the target version range.
[0071] In this embodiment, first screen from the preset vulnerability database based on the corrected version range, and then determine the target vulnerability test script. That is to say, according to the target software version range defined by the corrected version range, search for vulnerabilities related to the software versions within this range in the database and extract the corresponding test scripts. For example, if the corrected version range is Apache 2.4.49 - 2.4.50, then search for all vulnerabilities affecting the Apache software within this version range in the vulnerability database and obtain the test scripts for these vulnerabilities. These scripts may include specific codes and operation steps for verifying the existence of vulnerabilities.
[0072] Furthermore, determine the attack payload based on the target vulnerability test script, and send the attack payload to the target system to obtain the system response result. Among them, the attack payload is a specific set of data or instructions generated according to the requirements of the test script, which is used to simulate the operation of an attacker exploiting a vulnerability in the target system. For example, for an Apache path traversal vulnerability test script, the attack payload may be to construct a special HTTP request containing malicious path parameters and attempt to access files or directories that should not be publicly available in the system. After sending the attack payload to the target system, the system will process these requests and return the response result. This response result contains various information returned by the server, such as the HTTP response code, response header content, response body data, etc. These information are the key basis for determining whether there are vulnerabilities in the target system.
[0073] Then, the system response result is matched with the preset vulnerability features to obtain a matching result. The preset vulnerability features are predefined criteria or patterns for determining whether a target system has specific vulnerabilities. These features are usually summarized based on the analysis of vulnerability principles and actual test data, and may include specific error prompt messages, returned file content formats, response code combinations, etc. For example, for the above Apache path traversal vulnerability, the preset vulnerability features may be that the content of the system's sensitive files appears in the response body, or a specific format of error information is returned. The system will compare and analyze the obtained system response result with these preset vulnerability features one by one. If the response result contains the content described by the preset vulnerability features, it is considered a successful match; otherwise, it is a failed match, thus obtaining an accurate matching result.
[0074] In a specific implementation, if the matching result indicates that there is a vulnerability in the target system, the scope of influence of the vulnerability in the target system is further determined. Different vulnerabilities may affect different version ranges of the software. By analyzing the vulnerability-related information and combining the actual situation of the target system, the specific software version range affected by the vulnerability is determined. For example, a certain Apache vulnerability may be clearly recorded to affect versions 2.4.40 - 2.4.49, and the corrected version range of the current target system is 2.4.49 - 2.4.50. After analysis, it is found that the vulnerability is still valid in version 2.4.49, then the scope of influence of this vulnerability intersects with the current version range.
[0075] Finally, based on the scope of influence of the vulnerability, the corrected version range is adjusted to obtain the target version range. That is, according to the relationship between the scope of influence of the vulnerability and the corrected version range, the version range is correspondingly narrowed or corrected. If the scope of influence of the vulnerability is completely included in the corrected version range, the version range is adjusted to be the same as the scope of influence of the vulnerability; if the scope of influence of the vulnerability only partially covers the corrected version range, the intersection of the two is taken as the new version range. For example, if the corrected version range is 2.4.49 - 2.4.50 and the scope of influence of the vulnerability is 2.4.40 - 2.4.49, the adjusted target version range is 2.4.49. Through such adjustment, the version range more accurately reflects the actual situation of the vulnerabilities existing in the target system.
[0076] Step S14: Determine the probability of the existence of a vulnerability in the target system based on the version probability distribution and the target version range.
[0077] In this embodiment, first, target vulnerabilities are searched from a preset vulnerability database based on a target version range. The target version range defines the possible version range of the target software, and retrieval is performed in the vulnerability database according to this range. For example, when the target version range is determined to be Apache 2.4.49, then all records in the database will be traversed, and vulnerability entries affecting the Apache 2.4.49 version will be screened out. These entries contain key information such as vulnerability numbers (e.g., CVE-2021-41773), vulnerability descriptions, and affected ranges, forming a set of target vulnerabilities related to the target system.
[0078] Then, in this embodiment, the version coverage range of the target vulnerability is determined, and the probability of the existence of the vulnerability in the target system is determined based on the version probability distribution and the version coverage range. Each target vulnerability records in the database the version range it affects. Taking the CVE-2021-41773 vulnerability that covers the 2.4.49 version with a probability of 65% as an example, then the probability of the existence of this vulnerability directly takes this version probability, that is, 65%. If a vulnerability such as CVE-2021-40438 affects v2.4.40 - 2.4.49, and the candidate versions are {v2.4.49: 65%, v2.4.50: 35%}, due to the relationship between the affected range and the version probability distribution, the probability of the existence of the vulnerability is 65%. When a vulnerability such as CVE-2021-42013 affects versions 2.4.49 and 2.4.50, since these two versions in the candidate versions are completely covered, the probability of the existence of the vulnerability is 100%. For other similar vulnerabilities, if their affected versions cover 2.4.49 and 2.4.50, the vulnerability impact probability is also 100%; while for vulnerabilities whose affected versions are less than 2.4.49, since they are not within the main range of the current version probability distribution, the vulnerability probability is 0%. Through this quantitative calculation method, the version information is deeply combined with the vulnerability affected range to achieve an accurate assessment of the possibility of the existence of the vulnerability.
[0079] Furthermore, the preset Common Vulnerability Scoring System (CVSS) is used to score the target vulnerability to obtain a scoring result. For example, for a high-risk vulnerability that can be exploited remotely and does not require special permissions, its basic score may reach 9.8 points; if there is publicly available exploit code for this vulnerability and it has not been fixed, the time score will be further increased. The system inputs corresponding parameters into the CVSS scoring model according to the specific attributes of the vulnerability and automatically generates a standardized scoring result, which intuitively reflects the severity of the vulnerability.
[0080] Meanwhile, determine the exploit status of the target vulnerability. The exploit status is used to characterize whether there is public or private exploit code (i.e., EXP, Exploit) for the target vulnerability, and this information can be obtained from channels such as security community reports and vulnerability intelligence platforms. For example, for some popular vulnerabilities, a large number of public exploit codes appear in a short period of time, and their exploit status is "public EXP exists"; while for some vulnerabilities that have not been widely studied, they may be in the state of "no public EXP". Clarifying the exploit status can further evaluate the risk of the vulnerability being actually attacked.
[0081] Finally, generate the corresponding risk level based on the scoring result and the exploit status, and use the risk level to perform the corresponding preset risk response operations. For example, a set of mapping rules can be established to convert the scoring and exploitation status into risk levels: when the CVSS score is greater than or equal to 9.0 and there is public EXP, it is determined as a "high-risk" risk; when the score is between 7.0 - 8.9 and there is EXP, or the score is greater than or equal to 9.0 but there is no public EXP, it is determined as "medium-risk"; when the score is less than 7.0 and there is no public EXP, it is determined as "low-risk". For different risk levels, corresponding response strategies are preset: for high-risk vulnerabilities, immediately trigger the emergency repair process, notify the operation and maintenance personnel to suspend the relevant services and deploy patches; for medium-risk vulnerabilities, it is required to complete the evaluation and implementation of the repair plan within 48 hours; low-risk vulnerabilities can be included in the regular inspection plan and processed during the next system maintenance. Through this hierarchical response mechanism, the efficient management of vulnerability risks is achieved, and the impact of security threats on the target system is minimized.
[0082] As can be seen from the above, during the software vulnerability detection process of this application, first, with the help of port scanning tools and Web service identification tools, the target system is analyzed to obtain the initial software version range of the target software. On this basis, the initial software version range is further optimized and corrected according to the file change time of the target file and the version release time of the target software, and then a more accurate corrected version range is obtained. Next, a pre-trained model based on the Transformer architecture is trained using a preset vulnerability database. Through the training process, the model continuously learns and masters the correlation between software versions and various features, and finally forms the target large language model, and the target large language model is used to determine the version probability distribution of the target software. Then, based on the corrected version range, the appropriate target vulnerability test script is screened out in the preset vulnerability database. The target system is tested using the target vulnerability test script, and based on the test results, the corrected version range is dynamically adjusted again, thereby obtaining a more practical target version range. Finally, the actual probability of each vulnerability existing in the target system is determined through the version probability distribution and the target version range, providing certain data support for subsequent vulnerability risk assessment and handling. In this way, this application can improve the efficiency of vulnerability detection, quickly locate risks, and reduce system security risks.
[0083] The following combines Figure 2 the schematic diagram shown to specifically describe the technical solution of the embodiment of this application.
[0084] Specifically, the first is multi-source detection and fusion. The system parallelly calls professional tools such as Nmap and Wappalyzer to comprehensively detect the target system. As a powerful port scanning tool, Nmap collects data such as network protocol identifiers, response header information, and service types through network interaction with the target system. For example, when scanning a certain Web server, Nmap obtains the protocol identifier of "SSH-2.0-OpenSSH_8.2p1" and the content of "Server: Apache / 2.4.x" in the HTTP response header, thereby initially judging the software and version clues running on the server. Wappalyzer focuses on the Web service level and determines the target software information from another dimension by parsing page elements, script features, and resource references. After the system obtains the first and second software version ranges respectively, it takes the intersection to obtain the initial version range. For example, the initial version range of the Apache software is determined to be v2.4.30 - 2.4.50, effectively integrating multi-source data and narrowing the version range.
[0085] Next is the time feature anchoring step. The system collects the release times of each version of the target software from a preset vulnerability database (i.e., the public vulnerability database) and the official channels of the manufacturers, and constructs a detailed version timeline. At the same time, analyze the static files in the target system. Taking the / api / v2 / docs file as an example, send an HTTP request to obtain its Last-Modified time as 2021-09-16. Compare this time with the version timeline and find that the file has changed after version v2.4.49, thus determining the lower version limit as v2.4.49 and further refining the version range.
[0086] Subsequently, enter the large model inference stage. Input the feature information such as response headers and protocol fingerprints obtained from multi-source detection into a pre-trained model based on the Transformer architecture. The model has been trained with a large amount of "version-feature" data and can deeply explore the correlation between features and versions. At the same time, adjust the probability distribution in combination with version usage (i.e., version download volume) data, and if a certain version fixes a major vulnerability, increase its probability weight. Through model calculation and analysis, finally output the version probability distribution. For example, the probability of Apache software v2.4.49 is 65% and the probability of v2.4.50 is 35%, providing a quantitative basis for version judgment.
[0087] Then perform the POC dynamic correction step. Based on the previously determined version range, screen the POC test scripts of relevant vulnerabilities from the vulnerability database, such as executing the POC test of CVE-2021-41773. After sending a carefully constructed attack payload to the target system, match the system response result with the preset vulnerability features. When it is confirmed that the vulnerability exists, according to the scope of influence of the vulnerability, correct the original version range from v2.4.40 - 2.4.50 to v2.4.49, realizing another precise adjustment of the version range.
[0088] Finally, there is vulnerability probability and risk assessment. Based on the determined version v2.4.49, find all the vulnerabilities affected by this version in the vulnerability database. Since v2.4.49 is completely within the scope of influence of the relevant vulnerabilities, it is confirmed that the probability of the vulnerability existing is 100%, and its risk level is evaluated as high risk.
[0089] Correspondingly, as shown in Figure 3 the embodiments of the present application provide a software vulnerability detection device, including:
[0090] The first interval determination module 11 is configured to determine an initial software version interval of the target software from the target system by using a port scanning tool and a web service identification tool, and correct the initial software version interval based on the file change time of the target file and the version release time of the target software to obtain a corrected version interval; wherein, the target file is included in the target software.
[0091] The probability determination module 12 is configured to train a pre-trained model based on the Transformer architecture by using a preset vulnerability database, and determine a version probability distribution based on the trained target large language model.
[0092] The second interval determination module 13 is configured to determine a target vulnerability test script from the preset vulnerability database based on the corrected version interval, and test the target system based on the target vulnerability test script to adjust the corrected version interval by using the obtained test results to obtain a target version interval.
[0093] The risk determination module 14 is configured to determine the probability of the existence of vulnerabilities in the target system based on the version probability distribution and the target version interval.
[0094] As can be seen from the above, in the process of software vulnerability detection in this application, first, with the help of a port scanning tool and a web service identification tool, the target system is analyzed to obtain the initial software version interval of the target software. On this basis, the initial software version interval is further optimized and corrected according to the file change time of the target file and the version release time of the target software, so as to obtain a more accurate corrected version interval. Next, a pre-trained model based on the Transformer architecture is trained by using a preset vulnerability database. Through the training process, the model continuously learns and masters the correlation between software versions and various features, and finally forms a target large language model, and the version probability distribution of the target software is determined by using the target large language model. Then, based on the corrected version interval, an appropriate target vulnerability test script is selected from the preset vulnerability database. The target system is tested by using the target vulnerability test script, and based on the test results, the corrected version interval is dynamically adjusted again, so as to obtain a more practical target version interval. Finally, the actual probability of the existence of each vulnerability in the target system is determined through the version probability distribution and the target version interval, providing certain data support for subsequent vulnerability risk assessment and handling. In this way, this application can improve the efficiency of vulnerability detection, quickly locate risks, and reduce system security risks.
[0095] In some specific embodiments, the first interval determination module 11 specifically includes:
[0096] A software information determination unit, which is used to determine the network protocol identifier, response header information, and service type of the target software from the target system by using a port scanning tool and a Web service identification tool respectively;
[0097] A first interval determination unit, which is used to determine a first software version interval based on the network protocol identifier, the response header information, and the service type through the port scanning tool;
[0098] A second interval determination unit, which is used to determine a second software version interval based on the network protocol identifier, the response header information, and the service type through the Web service identification tool;
[0099] An initial interval determination unit, which is used to determine the intersection interval of the first software version interval and the second software version interval, and determine the intersection interval as the initial software version interval.
[0100] In some specific embodiments, the first interval determination module 11 specifically includes:
[0101] A release time determination unit, which is used to determine the version release time of the target software from the preset vulnerability database and the preset manufacturer information;
[0102] A change time determination unit, which is used to determine a target static file from the target software, and use a preset hypertext transfer protocol request to determine the file change time of the target static file that is closest to the current time;
[0103] An interval correction unit, which is used to determine a version timeline based on the version release time and the file change time, and use the version timeline to correct the initial software version interval to obtain a corrected version interval.
[0104] In some specific embodiments, the probability determination module 12 specifically includes:
[0105] A model configuration unit, which is used to configure a multi-source feature encoding layer for the input layer of a pre-trained model based on the Transformer architecture, and configure a normalized exponential function for the output layer of the pre-trained model to obtain a configured model;
[0106] A model training unit, which is used to obtain training data from the preset vulnerability database and the preset manufacturer information, and train the configured model based on the training data to obtain a target large language model;
[0107] A first information determination unit, configured to determine the website request return information about the target software and the time information of the target static file obtained from the target system by using the port scanning tool and the Web service identification tool, so as to determine first input information;
[0108] A second information determination unit, configured to determine the version usage data and version priority of the target software as second input information;
[0109] An information input unit, configured to input the first input information and the second input information into the target large language model, so that the target large language model outputs a version probability distribution.
[0110] In some specific embodiments, the second interval determination module 13 specifically includes:
[0111] A payload sending unit, configured to determine an attack payload based on the target vulnerability test script and send the attack payload to the target system to obtain a system response result;
[0112] A result matching unit, configured to match the system response result with a preset vulnerability feature to obtain a matching result;
[0113] A scope determination unit, configured to determine the vulnerability impact scope in the target system if the matching result indicates that there is a vulnerability in the target system;
[0114] A target interval determination unit, configured to adjust the corrected version interval based on the vulnerability impact scope to obtain a target version interval.
[0115] In some specific embodiments, the risk determination module 14 specifically includes:
[0116] A vulnerability search unit, configured to search for a target vulnerability from the preset vulnerability database based on the target version interval;
[0117] A probability determination unit, configured to determine the version coverage range of the target vulnerability and determine the vulnerability existence probability in the target system based on the version probability distribution and the version coverage range.
[0118] In some specific embodiments, the risk determination module 14 specifically further includes:
[0119] A vulnerability scoring unit, configured to score the target vulnerability by using a preset common vulnerability scoring system to obtain a scoring result;
[0120] A status determination unit, configured to determine the vulnerability exploitation status of the target vulnerability; the vulnerability exploitation status indicates whether there is vulnerability exploitation code for the target vulnerability.
[0121] A risk response unit, configured to generate a corresponding risk level based on the scoring result and the vulnerability exploitation status, and perform a corresponding preset risk response operation using the risk level.
[0122] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the software vulnerability detection method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0123] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0124] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.
[0125] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the software vulnerability detection method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.
[0126] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the software vulnerability detection method disclosed above. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not repeated here.
[0127] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and for the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method section.
[0128] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0129] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0130] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0131] The above has introduced the technical solution provided by this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A software vulnerability detection method, characterized in that Including: Using a port scanning tool and a web service identification tool to determine an initial software version range of the target software from the target system, and correcting the initial software version range based on the file change time of the target file and the version release time of the target software to obtain a corrected version range; wherein, the target file is included in the target software; Training a pre-trained model based on the Transformer architecture using a preset vulnerability database, and determining a version probability distribution based on the trained target large language model; Determining a target vulnerability test script from the preset vulnerability database based on the corrected version range, and testing the target system based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range; Determining the probability of vulnerability existence in the target system based on the version probability distribution and the target version range.
2. The software vulnerability detection method according to claim 1, wherein The step of using a port scanning tool and a web service identification tool to determine an initial software version range of the target software from the target system includes: Respectively using a port scanning tool and a web service identification tool to determine the network protocol identifier, response header information, and service type of the target software from the target system; Determining a first software version range through the port scanning tool based on the network protocol identifier, the response header information, and the service type; Determining a second software version range through the web service identification tool based on the network protocol identifier, the response header information, and the service type; Determining the intersection range of the first software version range and the second software version range, and determining the intersection range as the initial software version range.
3. The software vulnerability detection method according to claim 1, characterized in that The step of correcting the initial software version range based on the file change time of the target file and the version release time of the target software to obtain a corrected version range includes: Determining the version release time of the target software from the preset vulnerability database and the preset manufacturer information; Determining a target static file from the target software, and using a preset hypertext transfer protocol request to determine the file change time of the target static file closest to the current time; Determining a version timeline based on the version release time and the file change time, and correcting the initial software version range using the version timeline to obtain a corrected version range.
4. The software vulnerability detection method according to claim 3, wherein The step of training a pre-trained model based on the Transformer architecture using a preset vulnerability database, and determining a version probability distribution based on the trained target large language model includes: Configuring a multi-source feature encoding layer for the input layer of the pre-trained model based on the Transformer architecture, and configuring a normalized exponential function for the output layer of the pre-trained model to obtain a configured model; Obtaining training data from the preset vulnerability database and the preset manufacturer information, and training the configured model based on the training data to obtain a target large language model; Determine the website request return information about the target software and the time information of the target static file obtained from the target system using the port scanning tool and the Web service identification tool to determine the first input information; Determine the version usage data and version priority of the target software as the second input information; Input the first input information and the second input information into the target large language model so that the target large language model outputs a version probability distribution.
5. The software vulnerability detection method according to claim 1, wherein Testing the target system based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range, including: Determine an attack payload based on the target vulnerability test script and send the attack payload to the target system to obtain a system response result; Match the system response result with a preset vulnerability feature to obtain a matching result; If the matching result indicates that there is a vulnerability in the target system, determine the scope of influence of the vulnerability in the target system; Adjust the corrected version range based on the scope of influence of the vulnerability to obtain a target version range.
6. The software vulnerability detection method according to any one of claims 1 to 5, characterized in that, The determining the probability of the existence of a vulnerability in the target system based on the version probability distribution and the target version range includes: Search for target vulnerabilities from the preset vulnerability database based on the target version range; Determine the version coverage range of the target vulnerability, and determine the probability of the existence of a vulnerability in the target system based on the version probability distribution and the version coverage range.
7. The software vulnerability detection method according to claim 6, wherein After determining the probability of the existence of a vulnerability in the target system based on the version probability distribution and the target version range, it further includes: Score the target vulnerability using a preset common vulnerability scoring system to obtain a scoring result; Determine the vulnerability exploitation status of the target vulnerability; the vulnerability exploitation status indicates whether there is vulnerability exploitation code for the target vulnerability; Generate a corresponding risk level based on the scoring result and the vulnerability exploitation status, and perform a corresponding preset risk response operation using the risk level.
8. A software vulnerability detection device, characterized in that, Including: A first interval determination module, configured to determine an initial software version range of the target software from the target system using a port scanning tool and a Web service identification tool, and correct the initial software version range based on the file change time of the target file and the version release time of the target software to obtain a corrected version range; wherein, the target file is included in the target software; A probability determination module, configured to train a pre-trained model based on the Transformer architecture using a preset vulnerability database, and determine a version probability distribution based on the trained target large language model; A second interval determination module, configured to determine a target vulnerability test script from the preset vulnerability database based on the corrected version range, and test the target system based on the target vulnerability test script to adjust the corrected version range using the obtained test results to obtain a target version range; A risk determination module, configured to determine the probability of the existence of vulnerabilities in the target system based on the version probability distribution and the target version interval.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the software vulnerability detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the software vulnerability detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Vulnerability influence range detection method and device, storage medium and electronic equipment
CN114996720A
Vulnerability influence range correction method and device, storage medium and electronic equipment
CN115455420A
Vulnerability repair information retrieval method and electronic equipment
CN115510446A
LLM-based ASOC vulnerability assessment method, apparatus and device, and medium
CN118395457A
Version Checking Apparatus, Version Checking System, and Version Checking Method
US20220179637A1