A Web application vulnerability detection method, device and storage medium
By determining the search status and clustering method based on the number of links and page frequency of the web application, more efficient and accurate vulnerability detection is achieved in response to the problems of excessive page crawling and single clustering methods in web application vulnerability detection.
Patent Information
- Application Number
- CN202411533242.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-31
AI Technical Summary
In vulnerability detection of existing web applications, excessive page crawling leads to poor efficiency and a single clustering method, resulting in poor vulnerability detection effect.
By obtaining web application information, determining the search status based on the average value of the number of links and the maximum link difference, selecting the appropriate search method (refer to threshold analysis or traversing page analysis), and determining the clustering method based on the page update frequency and subpage frequency, including clustering of page differences, image similarity and comprehensive similarity, obtaining similar page combinations, and vulnerability detection is performed for each group.
It improves the efficiency and accuracy of vulnerability detection, avoids the problems of repeated detection and inaccurate clustering, and enhances the detection ability of similar pages.
Smart Images

Figure CN119420525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a method, device, and storage medium for detecting Web application vulnerabilities. Background Art
[0002] In the existing detection of Web application vulnerabilities, there are problems such as excessive page crawling during page crawling, resulting in poor vulnerability detection efficiency, and a single clustering method for similar pages during clustering, resulting in poor vulnerability detection effects. Therefore, how to select pages and cluster similar pages to improve vulnerability detection efficiency is a technical problem that needs to be solved urgently by those skilled in the art.
[0003] Chinese Patent Publication No. CN112507341A discloses a vulnerability scanning method, device, equipment, and storage medium based on a web crawler, including: obtaining corresponding web pages to be crawled and the valid URLs of the web pages to be crawled according to a preset list of pages to be crawled; sorting and filtering the valid URLs, and updating the current crawler depth; converting the valid URLs containing preset keywords into standard URLs; checking for duplicates of the standard URLs in a preset crawling set, and updating the crawling set according to the results of the duplicate check; when the list of pages to be crawled is empty or the crawler depth exceeds a preset depth limit value, terminating the crawler scan. It can be seen that the above technical solution has the following problems: only obtaining URLs based on the list of pages to be crawled, unable to change the URL acquisition method according to the actual scenario, with a single URL selection method, and not clustering similar pages, and repeatedly detecting the URLs of similar pages, resulting in poor vulnerability detection efficiency. Summary of the Invention
[0004] For this reason, the present invention provides a method for detecting Web application vulnerabilities to overcome the problems in the prior art that only obtaining URLs based on the list of pages to be crawled, unable to change the URL acquisition method according to the actual scenario, with a single URL selection method, and not clustering similar pages, and repeatedly detecting the URLs of similar pages, resulting in poor vulnerability detection efficiency.
[0005] To achieve the above object, the present invention provides a method for detecting Web application vulnerabilities, including:
[0006] Obtaining Web application information, and determining the search state according to the average number of links and the maximum link difference;
[0007] Determining the search method as reference threshold analysis or traversing page analysis according to the search state;
[0008] Determining the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value;
[0009] Determine the clustering method according to the page type, which is to cluster URLs according to page difference degree or perform primary clustering according to image similarity and secondary clustering according to comprehensive similarity to obtain several combinations of similar pages;
[0010] For each combination of similar pages, extract a URL for vulnerability detection. After inputting the attack vector, determine whether there is a vulnerability according to the execution situation of the Web application.
[0011] Furthermore, when the search status is that the average number of links is greater than or equal to the preset average number of links or the maximum link difference is greater than or equal to the preset maximum link difference, the search method is reference threshold analysis, and the reference threshold analysis includes:
[0012] Perform reference threshold analysis on each initial sub-page corresponding to the initial page. When performing reference threshold analysis on a single initial sub-page, record the initial sub-page as the target sub-page, detect the reference threshold corresponding to the target sub-page. If the reference threshold corresponding to the target sub-page is greater than the preset reference threshold, record the target sub-page as the sub-page to be detected, and record each sub-page to be detected and the initial page as the pages to be detected;
[0013] The reference threshold is determined according to the sum of the page threshold corresponding to the target sub-page and the page thresholds corresponding to each associated sub-page.
[0014] Furthermore, when the search status is that the average number of links is less than the preset average number of links and the maximum link difference is less than the preset maximum link difference, the search method is traversal page analysis, and the traversal page analysis includes:
[0015] Record the initial page and each initial sub-page corresponding to the initial page as the pages to be detected.
[0016] Furthermore, determine the page type of the pages to be detected according to the page update frequency and the sub-page frequency reference value. The page types include:
[0017] One type of page with a page update frequency less than the preset page update frequency and a sub-page frequency reference value less than the preset sub-page frequency reference value;
[0018] Two types of pages with a page update frequency greater than or equal to the preset page update frequency or a sub-page frequency reference value greater than or equal to the preset sub-page frequency reference value.
[0019] Furthermore, determine the clustering method according to the page type;
[0020] For one type of page, cluster URLs according to page difference degree;
[0021] For two types of pages, perform primary clustering according to image similarity and secondary clustering according to comprehensive similarity.
[0022] Further, URL clustering is performed for each first-class page according to the page difference degree, including:
[0023] Performing similar page detection for each first-class page corresponding to each URL. When performing similar page detection for a first-class page corresponding to a single URL, the first-class page corresponding to this URL is denoted as the target page, and the first-class pages corresponding to other URLs are denoted as reference pages. Detect the page difference degree between the target page and each reference page, and denote the set of all reference pages and the target page whose page difference degree from the target page is less than the preset page difference degree as a similar page combination, and continue to perform similar page detection for the reference pages that have not been denoted as similar page combinations until all first-class pages corresponding to all URLs are denoted as similar page combinations.
[0024] Further, perform primary clustering for each second-class page according to the image difference degree to obtain a number of primary clustering combinations;
[0025] The image difference degree is determined according to the feature region difference value and the site mean difference value;
[0026] The relationship between the feature region difference value and the image difference degree is a positive correlation relationship, and the relationship between the site mean difference value and the image difference degree is a positive correlation relationship.
[0027] Further, perform secondary clustering for each primary clustering combination according to the comprehensive similarity degree to obtain a number of similar page combinations:
[0028] The comprehensive similarity degree is determined according to the apparent similarity degree and the structure difference degree;
[0029] The relationship between the apparent similarity degree and the comprehensive similarity degree is a positive correlation relationship, and the relationship between the structure difference degree and the comprehensive similarity degree is a negative correlation relationship.
[0030] The present invention also provides a Web application vulnerability detection device, including:
[0031] A status analysis module, used to obtain Web application information, determine the search status according to the average value of the number of links and the maximum link difference, and determine the search method as reference threshold analysis or traversing page analysis according to the search status;
[0032] A page clustering module, which is connected to the status analysis module, used to determine the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value, and determine the clustering method as URL clustering according to the page difference degree or primary clustering according to the image similarity degree and secondary clustering according to the comprehensive similarity degree to obtain a number of similar page combinations;
[0033] A vulnerability detection module, which is connected to the page clustering module, is used to extract a URL for vulnerability detection for each combination of similar pages. After inputting an attack vector, it determines whether there is a vulnerability according to the execution situation of the Web application program.
[0034] The present invention also provides a storage medium, in which a computer program is stored. When the computer program runs on a computer, it can execute the Web application vulnerability detection method.
[0035] Compared with the prior art, the beneficial effect of the present invention is that in the technical solution of the present invention, the search state is determined according to the average number of links and the maximum link difference. The average number of links and the maximum link difference are used to effectively predict the number of URLs to be searched. Then, different search methods are adaptively selected according to the search state, making the selection of the search method more in line with the actual application scenario. It avoids the problem in the prior art that only based on the list to be crawled to obtain URLs and cannot change the URL acquisition method according to the actual scenario, and also avoids the problem of poor search effect caused by excessive page crawling during page crawling, improves the efficiency of searching for URLs, and thus improves the vulnerability detection efficiency.
[0036] Furthermore, in the present invention, the page type of the page to be detected is determined according to the page update frequency and the sub-page frequency reference value. The page update frequency and the sub-page frequency reference value effectively reflect the update situation of the page to be detected. Then, different clustering methods are adaptively selected according to the page type, making the selection of the clustering method more in line with the actual application scenario. It avoids the problem in the prior art that similar pages cannot be clustered, resulting in the inability to discover more vulnerabilities when detecting vulnerabilities for similar pages. It can perform a single detection on the URLs of similar pages, which is beneficial to improving the vulnerability detection efficiency, and also avoids the problem of inaccurate vulnerability detection caused by a single clustering method, and thus improves the accuracy of vulnerability detection.
[0037] Furthermore, in the present invention, reference threshold analysis is performed, which can filter URLs with fewer new pages when the number of URLs is large, avoiding the problem of poor page crawling rate, improving the efficiency of searching for URLs, and thus improving the vulnerability detection efficiency. Traversing page analysis is performed, which can ensure the page crawling rate when the number of URLs is small, avoiding the problem of inaccurate vulnerability detection, and thus improving the accuracy of vulnerability detection.
[0038] Furthermore, in the present invention, the page difference degree effectively reflects the difference situation of a class of pages, and then URL clustering is performed according to the page difference degree, avoiding the problem in the prior art that clustering cannot be performed for similar pages, resulting in the inability to discover more vulnerabilities when detecting vulnerabilities for similar pages, and also avoiding the problem of poor accuracy of the similarity of each page. Then, URL clustering is performed according to the page difference degree, improving the vulnerability detection efficiency.
[0039] Furthermore, in the present invention, primary clustering is performed according to the image similarity, which can preliminarily classify the pages that may have potential connections, thereby improving the similarity of each secondary page within the combined similar pages and avoiding the problem of poor clustering accuracy. Secondary clustering is performed according to the comprehensive similarity, which can more accurately classify the pages, avoiding the problem in the prior art that clustering cannot be performed for similar pages, resulting in the inability to discover more vulnerabilities when detecting vulnerabilities for similar pages. At the same time of clustering, the accuracy of clustering can be improved, thereby improving the vulnerability detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic diagram of the Web application vulnerability detection method of the present invention;
[0041] Figure 2 is a flowchart of determining the search method according to the search state of the present invention;
[0042] Figure 3 is a flowchart of the page type determination clustering method of the present invention;
[0043] Figure 4 is a schematic diagram of the Web application vulnerability detection device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to make the objectives and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0046] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.
[0047] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0048] Please refer to Figures 1 to 3 As shown, the present invention provides a method for detecting Web application vulnerabilities, including:
[0049] Obtain Web application information, and determine the search status according to the average number of links and the maximum link difference;
[0050] Determine the search method as reference threshold analysis or traversal page analysis according to the search status;
[0051] Determine the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value;
[0052] Determine the clustering method according to the page type as URL clustering according to page difference or primary clustering according to image similarity and secondary clustering according to comprehensive similarity to obtain several similar page combinations;
[0053] For each similar page combination, extract a URL for vulnerability detection. After inputting the attack vector, determine whether there is a vulnerability according to the execution situation of the Web application.
[0054] In the present invention, the Web application information includes but is not limited to the initial page, URL, and page link. The URL is a uniform resource locator. Each page corresponds to a URL. A single page contains several links, and each link corresponds to a sub-page. The initial page is the first page accessed when the web crawler performs data scraping operations. After performing a preset number of scraping operations on the URL of the initial page, the pages corresponding to the obtained URLs are recorded as screened pages. This is content that is easily understood by those skilled in the art and will not be elaborated here.
[0055] The value of the preset number can be determined by the user according to the actual application scenario. The higher the user's demand for improving the prediction accuracy of the URL quantity, the larger the value of the preset number. A value of the preset number is provided, and the prediction number is 5.
[0056] The search status includes: the average number of links is greater than or equal to the preset average number of links or the maximum link difference is greater than or equal to the preset maximum link difference, and the average number of links is less than the preset average number of links and the maximum link difference is less than the preset maximum link difference;
[0057] The average number of links is the average of the number of links corresponding to each filtered page. The method for determining the maximum link difference is to detect the number of links corresponding to each filtered page, and record the difference between the maximum value and the minimum value of the number of links as the maximum link difference; the number of links is the number of links included in a single filtered page.
[0058] For the values of the preset average number of links and the preset maximum link difference, the user can determine them according to the actual application scenario. The greater the user's demand for improving the search URL efficiency, the smaller the values of the preset average number of links and the preset maximum link difference. Provide a way to determine the values of the preset average number of links and the preset maximum link difference. Detect each historical vulnerability detection record, and record the average of the average number of links corresponding to the records that can meet the user's search URL efficiency requirements as the preset average number of links, and record the average of the maximum link differences corresponding to the records that can meet the user's search URL efficiency requirements as the preset maximum link difference.
[0059] Specifically, when the search status is that the average number of links is greater than or equal to the preset average number of links or the maximum link difference is greater than or equal to the preset maximum link difference, the search method is reference threshold analysis. The reference threshold analysis includes:
[0060] Perform reference threshold analysis on each initial sub-page corresponding to the initial page. When performing reference threshold analysis on a single initial sub-page, record the initial sub-page as the target sub-page, detect the reference threshold corresponding to the target sub-page. If the reference threshold corresponding to the target sub-page is greater than the preset reference threshold, record the target sub-page as the sub-page to be detected, and record each sub-page to be detected and the initial page as the pages to be detected;
[0061] The reference threshold is determined according to the sum of the page threshold corresponding to the target sub-page and the page thresholds corresponding to each associated sub-page.
[0062] Among them, a tree structure diagram is established based on the initial page. Node analysis is performed on the initial page to extract all sub-pages corresponding to the initial page. Several new sub-nodes are created for each sub-page, and each sub-node is connected to the root node. Node analysis is also performed on each sub-page, and each sub-node is connected to the corresponding sub-node until the page threshold of the sub-page is 0, at which point the node analysis stops. Finally, a tree structure diagram with the initial page as the root node and each sub-page as a sub-node is established. This is content that is easily understood by those skilled in the art and will not be elaborated here. The sub-pages corresponding to each sub-node in the tree structure diagram are denoted as initial sub-pages.
[0063] For the value of the preset reference threshold, the user can determine it according to the actual application scenario. The greater the user's value requirement for capturing sub-pages, the greater the value of the preset reference threshold. A value of the preset reference threshold is provided, and the preset reference threshold is 5.
[0064] The confirmation method for associated sub-pages is as follows: for a single initial sub-page, this initial sub-page is denoted as the target sub-page. The associated sub-pages of the target sub-page include the initial sub-pages corresponding to the nodes adjacent to the left of the node corresponding to the target sub-page in the tree structure diagram and the initial sub-pages corresponding to the nodes adjacent to the right of the node corresponding to the target sub-page in the tree structure diagram. It can be understood that if the target sub-page is located at the leftmost or rightmost position in the tree structure diagram, there is only one associated sub-page.
[0065] The page threshold is the number of links contained in a single sub-page.
[0066] Specifically, when the search status is that the average link number is less than the preset average link number and the maximum link difference is less than the preset maximum link difference, the search method is traversing page analysis, and the traversing page analysis includes:
[0067] The initial page and each initial sub-page corresponding to the initial page are denoted as pages to be detected.
[0068] Specifically, the page type of the page to be detected is determined according to the page update frequency and the sub-page frequency reference value. The page types include:
[0069] One type of page where the page update frequency is less than the preset page update frequency and the sub-page frequency reference value is less than the preset sub-page frequency reference value;
[0070] Two types of pages where the page update frequency is greater than or equal to the preset page update frequency or the sub-page frequency reference value is greater than or equal to the preset sub-page frequency reference value.
[0071] Among them, the page update frequency is the shortest time for the content of a single page to be detected to be updated, and the sub-page frequency reference value is the average value of the shortest times for the content of each sub-page corresponding to a single page to be detected to be updated.
[0072] For the values of the preset page update frequency and the reference value of the preset sub - page frequency, users can determine them according to the actual application scenario. The larger the values of the preset page update frequency and the reference value of the preset sub - page frequency, the smaller the degree of change of the page content. A method for obtaining the values of the preset page update frequency and the reference value of the preset sub - page frequency is provided. The preset page update frequency is 2h. The method for determining the reference value of the preset sub - page frequency is to detect the records with the same page update frequency as the current page to be detected during the historical clustering process, and take the average value of the sub - page frequency reference values of the records that can meet the user's vulnerability detection requirements as the reference value of the preset sub - page frequency.
[0073] Specifically, determine the clustering method according to the page type;
[0074] For one type of page, perform URL clustering according to the page difference degree;
[0075] For the second type of page, perform a first - stage clustering according to the image similarity and a second - stage clustering according to the comprehensive similarity.
[0076] Specifically, perform URL clustering for each first - type page according to the page difference degree, including:
[0077] Perform similar - page detection for the first - type pages corresponding to each URL. When performing similar - page detection for the first - type page corresponding to a single URL, record the first - type page corresponding to this URL as the target page, and record the first - type pages corresponding to other URLs as reference pages. Detect the page difference degree between the target page and each reference page. Denote the set of all reference pages and the target page whose page difference degree from the target page is less than the preset page difference degree as a similar - page combination, and continue to perform similar - page detection for the reference pages that have not been denoted as similar - page combinations until all the first - type pages corresponding to all URLs are denoted as similar - page combinations.
[0078] Among them, for two first - type pages, the method for determining the page difference degree is to use the hash algorithm to convert the text content corresponding to the first - type page into a 64 - bit binary string. This is understandable to those skilled in the art and will not be elaborated here. Then compare the two binary strings bit by bit, and denote the number of different bits as the page difference degree.
[0079] For the value of the preset page difference degree, users can determine it according to the actual application scenario. It can be understood that the smaller the value of the preset page difference degree, the greater the similarity of each first - type page in the similar - page combination, and the greater the accuracy of the user's vulnerability detection. A method for obtaining the value of the preset page difference degree is provided. Detect the records of URL clustering during the historical vulnerability detection process, and take the average value of the page difference degrees of the records that can meet the user's vulnerability detection requirements as the preset page difference degree.
[0080] Specifically, perform a clustering on each secondary page according to the image difference degree to obtain several first-level clustering combinations;
[0081] The image difference degree is determined according to the feature region difference value and the site mean difference value;
[0082] The relationship between the feature region difference value and the image difference degree is a positive correlation relationship, and the relationship between the site mean difference value and the image difference degree is a positive correlation relationship.
[0083] Among them, performing a clustering on each secondary page according to the image difference degree to obtain several first-level clustering combinations includes:
[0084] Perform a first-level clustering detection on each secondary page. When performing a first-level clustering detection on a single secondary page, mark this secondary page as the target secondary page, and mark the other secondary pages excluding the target secondary page as the reference secondary pages. Detect the image difference degree between the target secondary page and each reference secondary page. Mark all the reference secondary pages and the target secondary page with an image difference degree less than the preset image difference degree as a first-level clustering combination, and continue to perform a first-level clustering detection on the reference secondary pages that have not been marked as first-level clustering combinations until all secondary pages are marked as first-level clustering combinations.
[0085] For the value of the preset image difference degree, the user can determine it according to the actual application scenario. The smaller the value of the preset image difference degree, the greater the similarity of each secondary page in the first-level clustering combination, and the greater the accuracy of the user vulnerability detection. Provide a value of the preset image difference degree, detect the records of the first-level clustering during the historical vulnerability detection process, and mark the average value of the image difference degrees corresponding to the records that can meet the user vulnerability detection requirements as the preset image difference degree.
[0086] Image difference degree = (feature region difference value / preset feature region difference value) × difference coefficient + (site mean difference value / preset site mean difference value) × site coefficient;
[0087] For a single Class-II page, the method for confirming the feature region is to divide the region according to the page function area, and mark each divided region as a feature region. For a feature region in a page, the page function area corresponding to this feature region may be the same as or different from the page function areas corresponding to other feature regions of this page; mark the feature regions corresponding to the same page function area as feature regions of the same category. The number of feature regions of the same category corresponding to each page may be different. The page function area is the area with different functions in a single page, and the page function area includes a navigation area, a header area, a content display area, a sidebar, a bottom area, an interaction area, and a promotion area. The segmentation of the feature region can be achieved through machine vision and a deep learning network, which is easy for those skilled in the art to understand and will not be elaborated here.
[0088] The feature region difference value = the proportion coefficient of the same feature region + the difference coefficient. The relationship between the proportion coefficient of the same feature region and the proportion of the number of the same category is a positive correlation, and the relationship between the difference coefficient and the maximum feature region difference is a positive correlation; the method for confirming the proportion of the number of the same category is as follows: for two Class-II pages, mark the number of feature regions of the same category corresponding to one Class-I and Class-II page as A, and mark the number of feature regions of the same category corresponding to the other Class-II page as B. The proportion of the number of the same category = the number of the same category / the larger value of A and B. The number of the same category is the number of feature regions of the same category that both Class-II pages have.
[0089] For a Class-I and Class-II page, mark each feature region with the same page function area in the Class-II page as a feature region combination. Perform region difference detection on each feature region combination. When performing region difference detection on a feature region combination in a Class-I and Class-II page, mark this feature region combination as the target feature region combination, and mark the feature region combination with the same page function area as the target feature region combination in the other Class-II page as the reference region combination. If there is no reference region combination in the other Class-II page, the region difference is the number of feature regions corresponding to the target feature region combination. If there is a reference region combination in the other Class-II page, mark the absolute value of the difference between the number of feature regions corresponding to the target feature region combination and the number of feature regions corresponding to the reference region combination as the region difference. The method for confirming the maximum feature region difference is as follows: for two Class-II pages, mark the maximum value of the region differences corresponding to each region of the two Class-II pages as the maximum feature region difference.
[0090] It can be understood that each Class-II page is a rectangular region. The method for confirming the site mean difference value is as follows: for a single Class-II page, take the upper left corner of this Class-II page as the origin, take the straight line extending horizontally to the right passing through the origin as the x-axis, and take the straight line extending vertically downward perpendicular to the x-axis as the y-axis to establish a rectangular coordinate system.
[0091] For the first and second type pages, the average value of the coordinates of the central positions of the respective feature regions corresponding to the second type pages is denoted as the site average value. The method for confirming the site average value difference is as follows: for two second type pages, the site average value corresponding to the first and second type pages is denoted as (x1, y1), and the site average value corresponding to the other second type page is denoted as (x2, y2). The method for confirming the site average value difference z is as follows: the square of the difference between x1 and x2 is denoted as the first parameter, the square of the difference between y1 and y2 is denoted as the second parameter, and the square root of the sum of the first parameter and the second parameter is denoted as the site average value difference; for a single feature region, the central position is the center of the minimum circumscribed circle of the feature region.
[0092] The values of the preset feature region difference and the preset site average value difference can be determined by the user according to the actual application scenario. It can be understood that the smaller the values of the preset feature region difference and the preset site average value difference, the greater the similarity of the second type pages in a single clustering combination, and the greater the accuracy of the user vulnerability detection. Provide a set of values for the preset feature region difference and the preset site average value difference. Detect the records of a single clustering during the historical vulnerability detection process, and denote the average value of the feature region differences corresponding to the records that can meet the user vulnerability detection requirements as the preset feature region difference, and denote the average value of the site average value differences corresponding to the records that can meet the user vulnerability detection requirements as the preset site average value difference;
[0093] The values of the difference coefficient and the site coefficient can be obtained by the user through learning the historical records using a deep learning convolutional neural network. It can be understood that in the present invention, the difference degree of the second type pages is reflected by the image difference degree. The user can use deep learning through historical user data to obtain the influence of the feature region difference value and the site average value difference value on the difference degree of the second type pages respectively, and then correspondingly select the values of the difference coefficient and the site coefficient. Among them, the difference coefficient + the site coefficient = 1. Provide a set of values for the difference coefficient and the site coefficient, the difference coefficient is 0.5, and the site coefficient is 0.5.
[0094] Specifically, perform secondary clustering on each single clustering combination according to the comprehensive similarity to obtain several similar page combinations:
[0095] The comprehensive similarity is determined according to the apparent similarity and the structural difference degree;
[0096] The relationship between the apparent similarity and the comprehensive similarity is a positive correlation relationship, and the relationship between the structural difference degree and the comprehensive similarity is a negative correlation relationship.
[0097] Among them, performing secondary clustering on each single clustering combination according to the comprehensive similarity to obtain several similar page combinations includes:
[0098] Perform secondary clustering detection on each of the secondary pages corresponding to a single primary clustering combination. When performing secondary clustering detection on a single secondary page, mark this secondary page as the target clustering page, and mark the other secondary pages in the primary clustering combination that do not include the target clustering page as reference clustering pages. Detect the comprehensive similarity between the target clustering page and each reference clustering page, and mark all the reference clustering pages with a comprehensive similarity greater than the preset comprehensive similarity to the target clustering page and the target clustering page as a similar page combination, and continue to perform secondary clustering detection on the reference clustering pages that have not been marked as similar page combinations until all the secondary pages corresponding to the primary clustering combination are marked as similar page combinations.
[0099] The value of the preset comprehensive similarity can be determined by the user according to the actual application scenario. The larger the value of the preset comprehensive similarity, the greater the similarity between the secondary pages in the similar page combination, and the greater the accuracy of the user's vulnerability detection. Provide a value of the preset comprehensive similarity, detect the records of secondary clustering during the historical vulnerability detection process, and record the average value of the comprehensive similarity corresponding to the records that can meet the user's vulnerability detection requirements as the preset comprehensive similarity.
[0100] Comprehensive similarity = (apparent similarity / preset apparent similarity) × apparent coefficient - (structural difference degree / preset structural difference degree) × structural coefficient;
[0101] The present invention is provided with a continuously cyclic monitoring period. At the end of each monitoring period, the state of the feature area is determined once. The duration of the monitoring period can be set according to the user's needs. The greater the user's requirement for the monitoring accuracy of the feature area, the smaller the duration of the monitoring period. For a feature area, continuously monitor for a preset duration. If the feature area has not changed during the monitoring periods corresponding to the preset duration, mark the feature area as a fixed feature area. The value of the preset duration can be set according to the actual application scenario. The higher the user's accuracy requirement for the determination of the feature area state, the greater the value of the preset duration. Provide a value of the preset duration, and the preset duration is 30 monitoring periods;
[0102] Apparent similarity = text similarity coefficient - distribution difference coefficient. The relationship between the text similarity coefficient and the text similarity of the fixed feature area is a positive correlation relationship. The relationship between the distribution difference coefficient and the distribution difference degree is a positive correlation relationship;
[0103] For two secondary pages, mark a single fixed feature area in the first and secondary pages as the target fixed feature area, and mark the fixed feature area in the other secondary page with the smallest site difference from the target fixed feature area and the same page function area as the target fixed feature area as the reference fixed feature area, and mark the target fixed feature area and the reference fixed feature area as a detection combination;
[0104] For two Class-II pages, the text similarity of the fixed feature area is the average of the text similarities of each detection combination. The way to confirm the text similarity is to convert the text into a measurable numerical form, use the TF-IDF technology for text conversion, and use the cosine similarity to calculate the cosine value of the included angle between two text vectors to measure the similarity in their directions. The cosine value of the included angle between the two text vectors is recorded as the text similarity, and the value range of the text similarity is [0,1].
[0105] The distribution difference degree = area difference coefficient + site difference coefficient. The relationship between the area difference coefficient and the average area difference is a positive correlation, and the relationship between the site difference coefficient and the average site difference is a positive correlation;
[0106] For two Class-II pages, a single feature area in the Class-I and Class-II pages is denoted as the target area, and the feature area in the other Class-II page with the smallest site difference from the target area and the same page functional area as the target area is denoted as the reference area. The target area and the reference area are denoted as an area combination;
[0107] The average area difference is the average of the area difference values of each area combination. The way to confirm the area difference value is that for the feature areas corresponding to an area combination, the absolute value of the difference between the area of one feature area and the area of the other feature area is recorded as the area difference value. The site difference coefficient is the average of the corresponding site differences of each area combination. The way to confirm the site difference is that for two feature areas, the central position of one feature area is denoted as (X1, Y1), and the central position of the other feature area is denoted as (X2, Y2). The method to confirm the site difference Z is to square the difference between X1 and X2 as the first difference parameter, square the difference between Y1 and Y2 as the second difference parameter, and take the square root of the sum of the first difference parameter and the second difference parameter as the site difference;
[0108] The way to confirm the structural difference degree is that for two Class-II pages, the Class-II pages are converted into DOM tree structures, and the minimum number of operations required to convert the DOM tree structure corresponding to the Class-I and Class-II pages into the DOM tree structure corresponding to the other Class-II page is recorded as the structural difference degree. The operations include definition, addition, deletion, and replacement.
[0109] The values of the preset apparent similarity and the preset structural difference can be determined by the user according to the actual application scenario. When the value of the preset apparent similarity is larger and the value of the preset structural difference is smaller, the similarity of each secondary page in the similar page combination is larger, and the accuracy of user vulnerability detection is higher. A method is provided to determine the values of the preset apparent similarity and the preset structural difference. Record the secondary clustering during the historical vulnerability detection process, and take the average value of the apparent similarities of the records that can meet the user's vulnerability detection requirements as the preset apparent similarity, and take the average value of the structural differences of the records that can meet the user's vulnerability detection requirements as the preset structural difference.
[0110] The values of the apparent coefficient and the structural coefficient can be obtained by the user through learning the historical records using a deep learning convolutional neural network. It can be understood that the present invention reflects the similarity degree of secondary pages through the comprehensive similarity. The user can use deep learning through historical user data to obtain the influence of the apparent similarity and the structural difference on the similarity degree of secondary pages, and then correspondingly select the values of the apparent coefficient and the structural coefficient. Among them, the apparent coefficient + the structural coefficient = 1. A method is provided to determine the values of the apparent coefficient and the structural coefficient. The apparent coefficient is 0.4 and the structural coefficient is 0.6.
[0111] In the present invention, a URL is extracted for vulnerability detection for each similar page combination, and each extracted URT is used as a vulnerability detection target. An attack vector is input into each vulnerability detection target, and according to the execution situation of the Web application program, it is judged whether there is a vulnerability. If the Web application program executes an error, there is a vulnerability. If the Web application program executes correctly, there is no vulnerability. The attack vector is generated by constructing an attack vector factor library and based on rule constraints, which is easy for those skilled in the art to understand and will not be elaborated here.
[0112] Please refer to Figure 4 as shown in the figure, which is a schematic diagram of the Web application program vulnerability detection device of the present invention. The present invention also provides a Web application program vulnerability detection device, including:
[0113] A status analysis module for obtaining Web application program information, determining the search status according to the average value of the link quantity and the maximum link difference, and determining the search method as reference threshold analysis or traversing page analysis according to the search status;
[0114] A page clustering module connected to the status analysis module for determining the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value, and determining the clustering method as URL clustering according to the page difference or primary clustering according to the image similarity and secondary clustering according to the comprehensive similarity to obtain several similar page combinations;
[0115] A vulnerability detection module, which is connected to the page clustering module, is used to extract a URL for vulnerability detection for each group of similar pages. After inputting an attack vector, it determines whether there is a vulnerability according to the execution situation of the Web application program.
[0116] The present invention also provides a storage medium, in which a computer program is stored. When the computer program runs on a computer, it can execute the Web application program vulnerability detection method.
[0117] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0118] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A Web application vulnerability detection method, characterized in that: include: Obtain web application information and determine the search status based on the average number of links and the maximum link difference; Determine the search method as reference threshold analysis or traversal page analysis to obtain the page to be detected according to the search state; Determine the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value; The clustering method is determined according to the page type, which is to cluster the URLs according to the page difference or to cluster them once according to the image similarity and then cluster them again according to the comprehensive similarity to obtain several similar page combinations; For each similar page combination, a URL is extracted for vulnerability detection. After the attack vector is input, it is determined whether there is a vulnerability based on the execution of the Web application.
2. The Web application vulnerability detection method according to claim 1, characterized in that: When the search status is that the average value of the number of links is greater than or equal to the preset average value of the number of links or the maximum link difference is greater than or equal to the preset maximum link difference, the search method is reference threshold analysis, which includes: A reference threshold analysis is performed on each initial sub-page corresponding to the initial page. When a reference threshold analysis is performed on a single initial sub-page, the initial sub-page is recorded as a target sub-page, and a reference threshold corresponding to the target sub-page is detected. If the reference threshold corresponding to the target sub-page is greater than a preset reference threshold, the target sub-page is recorded as a sub-page to be detected, and each sub-page to be detected and the initial page are recorded as pages to be detected; The reference threshold is determined according to the sum of the page threshold corresponding to the target subpage and the page threshold corresponding to each associated subpage.
3. The Web application vulnerability detection method according to claim 2, characterized in that: When the search status is that the average link quantity is less than the preset link quantity average and the maximum link difference is less than the preset maximum link difference, the search method is traversal page analysis, which includes: The initial page and each initial sub-page corresponding to the initial page are recorded as pages to be detected.
4. The Web application vulnerability detection method according to claim 3, characterized in that: The page type of the page to be detected is determined according to the page update frequency and the sub-page frequency reference value. The page types include: A type of page whose page update frequency is less than a preset page update frequency and whose sub-page frequency reference value is less than a preset sub-page frequency reference value; The second category of pages has a page update frequency greater than or equal to a preset page update frequency or a sub-page frequency reference value greater than or equal to a preset sub-page frequency reference value.
5. The Web application vulnerability detection method according to claim 4, characterized in that: Determine the clustering method based on page type; For a type of page, URL clustering is performed based on page differences; For the second category pages, a clustering is performed based on image similarity, and a secondary clustering is performed based on comprehensive similarity.
6. The Web application vulnerability detection method according to claim 5, characterized in that: URL clustering is performed for each category of pages based on page differences, including: Similar page detection is performed on a type of page corresponding to each URL. When similar page detection is performed on a type of page corresponding to a single URL, the type of page corresponding to the URL is recorded as the target page, and the type of pages corresponding to other URLs are recorded as reference pages. The page difference between the target page and each reference page is detected, and the set of all reference pages and the target page whose page difference with the target page is less than a preset page difference is recorded as a similar page combination. Similar page detection is continued for reference pages that are not recorded as similar page combinations until all type of pages corresponding to all URLs are recorded as similar page combinations.
7. The Web application vulnerability detection method according to claim 6, characterized in that: Performing a clustering operation on each of the two types of pages according to the image difference to obtain a number of primary clustering combinations; The image difference is determined based on the feature area difference value and the site mean difference value; The relationship between the characteristic region difference value and the image difference is a positive correlation, and the relationship between the site mean difference value and the image difference is a positive correlation.
8. The Web application vulnerability detection method according to claim 7, characterized in that: Perform secondary clustering on each primary clustering combination according to the comprehensive similarity to obtain several similar page combinations: The comprehensive similarity is determined based on the apparent similarity and the structural difference; The relationship between the apparent similarity and the comprehensive similarity is a positive correlation, and the relationship between the structural difference and the comprehensive similarity is a negative correlation.
9. A Web application vulnerability detection device using the method described in any one of claims 1 to 8, characterized in that: include: A state analysis module is used to obtain Web application information, determine the search state according to the average value of the number of links and the maximum link difference, and determine the search method as reference threshold analysis or traversal page analysis according to the search state; A page clustering module, which is connected to the state analysis module, is used to determine the page type of the page to be detected according to the page update frequency and the sub-page frequency reference value, and to determine the clustering method according to the page type, which is to perform URL clustering according to page difference or to perform a primary clustering according to image similarity and a secondary clustering according to comprehensive similarity, to obtain a plurality of similar page combinations; The vulnerability detection module is connected to the page clustering module and is used to extract a URL for vulnerability detection for each similar page combination. After inputting the attack vector, it determines whether there is a vulnerability based on the execution status of the Web application.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program runs on a computer, the Web application vulnerability detection method according to any one of claims 1 to 8 can be executed.
Citation Information
Patent Citations
Loophole scanning method and device based on web crawler, equipment and storage medium
CN112507341A
Webpage recognition method and device
CN104462152A
Web page group extraction method, device and program
JP2010123000A