CMS identification method based on limited request packet detection and multi-feature detection
By constructing multiple scan content regions and feature regions based on limited request packet detection and multi-feature detection methods, the existing CMS recognition methods reduce scanning speed and inaccurate detection results when the feature library is enlarged, and more accurate and efficient CMS recognition is achieved.
Patent Information
- Application Number
- CN202510240054.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing CMS recognition methods lead to a decrease in scanning speed and inaccurate detection results when the feature library is enlarged, and the feature matching method has a high risk of underreport.
CMS recognition method based on limited request packet detection and multi-feature detection is adopted. By obtaining home page analysis, crawling content and resource files, calculating MD5 and MinHash, building multiple scan content regions and feature regions, and scanning matching recognition is performed within a limited range.
It effectively controls the scanning speed and the number of packets sent, reduces the risk of inaccurate detection results, improves the accuracy and detection rate of detection results, and reduces the probability of missed reports.
Abstract
Description
Technical Field
[0001] The invention discloses a CMS identification method based on limited request packet detection and multi-feature detection, which belongs to the technical field of network security. Background Art
[0002] Content Management System, or CMS for short. A content management system is a programming language that runs on the server to manage and maintain the columns, content, and templates of a website. System website CMS usually includes multiple CMS programs that implement different functions. Currently, CMS identification mainly relies on the fingerprint library, which stores a large number of CMS fingerprints. By extracting the CMS fingerprint of the web page to be identified and comparing the extracted CMS fingerprint with the fingerprint in the fingerprint library, the CMS category of the web page can be determined.
[0003] In the current network security industry, most methods for identifying CMS use single-feature or limited-feature detection technology, such as homepage source code, homepage header, static file content, static file MD5 and other features, and feature matching uses regular or identical expressions for detection. However, the existing detection methods still have the following problems:
[0004] 1. With the increase of feature records in the feature library, the number of detection data packets sent by the detection engine will also increase. When the feature value is too large, the scanning and recognition speed will be reduced. On the other hand, excessive packet sending will cause excessive pressure on the detected website and affect the normal operation of the business. At the same time, if the detected website is protected by security equipment, the probability of being intercepted by the security equipment will increase, resulting in inaccurate detection results.
[0005] 2. The feature equality or regular matching method has a high rate of false negatives and cannot effectively detect CMS systems that have been redeveloped or modified without features. Summary of the invention
[0006] In view of the shortcomings of the existing technology, in order to avoid the problems of reduced scanning speed, inaccurate detection results, and high vulnerability in feature matching as the feature library increases during the CMS identification process, a CMS identification method based on limited request packet detection and multi-feature detection is proposed.
[0007] The CMS identification method based on limited request packet detection and multi-feature detection is characterized by comprising the following steps:
[0008] S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page;
[0009] S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions;
[0010] S3, build a scan context, combining multiple scan content regions and feature regions in S2 to build a complete scan context;
[0011] S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.
[0012] Furthermore, in the step S1, the homepage content is the original data of the homepage obtained, and these data are usually in HTML format, including various elements and reference files required to construct the homepage;
[0013] Parse the acquired homepage content, identify and extract various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and picture tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files;
[0014] The image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.
[0015] Furthermore, in the step S2, the parsed HTML content is crawled, and multiple content regions and feature regions are constructed through content acquisition and feature calculation, which are specifically as follows:
[0016] (1) Extract HTML tags to form label regions;
[0017] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in HTML to form a filelist region;
[0018] (3) Download all files in the filelist region to form the file region;
[0019] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;
[0020] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;
[0021] (6) Extract the external link URL in HTML to form a link region;
[0022] Furthermore, in the step S3, a scanning context is constructed, which specifically includes: creating an initial scanning context framework, incorporating the multiple scanning content regions and feature regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple scanning content regions and feature regions incorporated, and then constructing a scanning context associated with the step S2.
[0023] Furthermore, in step S4, the rule recording method adopts a points system, and the match is successful if the specified points are met; wherein the rule record specifies the points required for the hit, which are specified by the @Match field in the record, and the @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained, and if there is no hit, no points are obtained. The points of all hit detection rules are added together, and if the total points are greater than or equal to the points of @Match4, the hit is recorded and an accurate scanning and recognition result is generated;
[0024] Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.
[0025] The detection rules specified by @Pattern include:
[0026] @Pattern 2region:label;
[0027] @Pattern 1region:file;
[0028] @Pattern 1region:filelist;
[0029] @Pattern 1region:md5;
[0030] @Pattern 1region:minihash;
[0031] @Pattern 1region:link;
[0032] Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.
[0033] The beneficial effects of the present invention are as follows: 1. First, by constructing multiple scanning content regions and feature regions through the parsed HTML content, the multiple scanning content regions and feature regions are then incorporated into the scanning context framework to establish their related detection environment, and to standardize the scanning range, thereby stipulating that the scanning is performed within a limited range. The possibility of reducing the scanning speed due to an increase in the number of acquired scanning elements will not occur as the scanning records increase. On the other hand, the number of data packets sent when scanning the scanning target is effectively controlled, thereby reducing the impact on the target, reducing the risk of inaccurate detection results, and effectively bypassing the target's protective measures.
[0034] 2. Construct multiple scanning content regions and feature regions, clarify the scanning scope, form a multi-feature and multi-region scanning matching method, thereby reducing the probability of missed reports and making the detection results more accurate. Compared with the existing technology, the single type of feature compared to the traditional scanning method improves the flexibility of scanning records and can further improve the detection rate. DETAILED DESCRIPTION
[0035] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.
[0036] This embodiment provides a technical solution: a CMS identification method based on limited request packet detection and multi-feature detection includes the following steps:
[0037] S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page;
[0038] The home page content is the original data of the home page, which is usually in HTML format and contains various elements and reference files required to build the home page.
[0039] Parse the obtained homepage content, identify and extract the various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and image tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files;
[0040] Image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.
[0041] Specifically, for https: / / www.example.comMonitoring detailed steps: send the scanning device to the desired scan through the URL http: / / www.example.com Request and obtain the home page content, usually in HTML format;
[0042] S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions;
[0043] Crawl the parsed HTML content and construct multiple scanning content regions and feature regions. Construct multiple content regions and feature regions through content acquisition and feature calculation. The details are as follows:
[0044] (1) Extract HTML tags to form label regions;
[0045] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in HTML to form a filelist region;
[0046] (3) Download all files in the filelist region to form the file region;
[0047] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;
[0048] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;
[0049] (6) Extract the external link URL in HTML to form a link region;
[0050] Specifically, the HTML content is parsed and multiple scanning content regions and feature regions are constructed, including:
[0051] (1) Extract HTML tags to form label regions;
[0052] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in html to form filelistregion;
[0053] (3) Download all files in the filelist region to form the file region;
[0054] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;
[0055] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;
[0056] (6) Extract the external link URLs in HTML to form a link region.
[0057] By constructing multiple scanning content regions and feature region combinations through a limited number of http detection packets, there is no need to send additional request packets during the scanning and matching feature process, which standardizes the scanning range during the CMS fingerprint recognition process and further stipulates that the scanning is performed within a limited range.
[0058] S3, build a scan context, combining multiple scan content regions and feature regions in S2 to build a complete scan context;
[0059] Constructing a scanning context specifically includes: creating an initial scanning context framework, incorporating the multiple scanning content regions and feature regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple scan regions included, and then constructing a scanning context associated with the S2 step.
[0060] Specifically, an initial scanning context framework is created, and the multiple scanning content regions and feature regions constructed in the above step S2, including label region, filelist region, file region, md5 region, minihash region and link region, are incorporated into the initial scanning context framework, so as to clarify that the scanning operation is specifically performed on the multiple scanning content regions and feature regions incorporated, and then a scanning context associated with step S2 is constructed, so that the scanning speed will not decrease with the increase of feature records in the feature library, thereby avoiding inaccurate detection results and the problem of underreporting.
[0061] S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.
[0062] The rule recording method adopts a points system. If the specified points are met, the match is successful. The rule record specifies the points required for the hit, which is specified by the @Match field in the record. The @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained. If there is no hit, no points are obtained. The points of all hit detection rules are added together. If the total points are greater than or equal to the points of @Match 4, the hit is recorded and an accurate scanning and recognition result is generated.
[0063] Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.
[0064] The detection rules specified by @Pattern include:
[0065] @Pattern 2region:label;
[0066] @Pattern 1region:file;
[0067] @Pattern 1region:filelist;
[0068] @Pattern 1region:md5;
[0069] @Pattern 1region:minihash;
[0070] @Pattern 1region:link;
[0071] Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.
[0072] The score required for a specified hit in the record is specified by the @Match field in the record. The @Pattern in the record specifies a detection rule. If the rule is hit, the specified score is obtained. If there is no hit, no score is obtained. The scores of all hit rules are added together. If the total score is greater than or equal to the score in @Match, the record is hit and an accurate scanning and recognition result is generated.
[0073] Specifically, the record example of the embodiment is:
[0074] @Match 4
[0075] @Pattern 2" "!region:lable;
[0076] @Pattern 1 "xxxs"! path"xxx.js"! region:file;
[0077] @Pattern 1"ruoyi.css"! region:filelist;
[0078] @Pattern 1"2221f431e845d608b9bb71c07c7a6a18"! path". / js / index.js"! region:md5;
[0079] @Pattern 1"11111"! path". / js / index.js"! Similarity 0.8! region:minihash;
[0080] @Pattern 1"http: / / xxx.com"! region:link;
[0081] Use the added CMS fingerprint library record file to scan and compare with the context constructed above. In the record, @Pattern2" "!region:label
[0082] Matches " in the html label tag "If you hit it, you get 2 points, if you miss it, you get no points;
[0083] @Pattern 1 "xxxs"! path"xxx.js"! region:file
[0084] Match "xxxs" in the xxx.js content in the file region and get 1 point;
[0085] @Pattern 1"ruoyi.css"! region:filelist
[0086] Check whether the filelist region contains the ruoyi.css file. If it does, you will get 1 point.
[0087] @Pattern 1"2221f431e845d608b9bb71c07c7a6a18"! path". / js / index.js"! region:md5
[0088] If the md5 of the file ". / js / index.js" is matched in the md5 region
[0089] "2221f431e845d608b9bb71c07c7a6a18" will get 1 point if hit;
[0090] @Pattern 1"11111"! path". / js / index.js"! Similarity 0.8! region:minihash
[0091] Match ". / js / index.js" in the minhash region to see if the similarity between the minihash value and "11111" is greater than or equal to 0.8. If so, a point is obtained.
[0092] @Pattern 1"http: / / xxx.com"! region:link
[0093] In the link region, check whether the external link "http: / / xxx.com" is included. If it is matched, a point is obtained. The points of all the matches are added together. If it is greater than or equal to the score specified in @Match, the rule is matched, and an accurate scanning and recognition result is output.
[0094] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A CMS identification method based on limited request packet detection and multi-feature detection, characterized in that: The following steps are involved: S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page; S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions; S3, build a scan context, combining multiple scan regions and feature regions in S2 to build a complete scan context; S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.
2. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S1, the home page content is the original data of the home page obtained, and this data is usually in HTML format, including various elements and reference files required to build the home page; Parse the acquired homepage content, identify and extract various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and picture tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files; The image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.
3. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S2, the parsed HTML content is crawled, and multiple content regions and feature regions are constructed through content acquisition and feature calculation, which are specifically as follows: (1) Extract HTML tags to form label regions; (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in html to form filelistregion; (3) Download all files in the filelist region to form the file region; (4) Calculate the md5 of the file content in the filelist region to form the md5 region; (5) Calculate the minihash of the file content in the filelist region to form a minihash region; (6) Extract the external link URLs in HTML to form a link region.
4. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S3, a scanning context is constructed, which specifically includes: creating an initial scanning context framework, incorporating the multiple scanning regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple included scanning regions, and on the basis of the included scanning regions, integrating the calculated features into the scanning context, and then constructing a scanning context associated with the step S2.
5. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In step S4, the rule recording method adopts a points system, and the match is successful if the specified points are met; the rule record specifies the points required for the hit, which is specified by the @Match field in the record, and the @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained, and if there is no hit, no points are obtained. The points of all hit detection rules are added together. If the total points are greater than or equal to the points of @Match 4, the hit is recorded and an accurate scanning and recognition result is generated; Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.
6. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 5 is characterized in that: The detection rules specified by @Pattern include: @Pattern 2region:label; @Pattern 1region:file; @Pattern 1region:filelist; @Pattern 1region:md5; @Pattern 1region:minihash; @Pattern 1region:link; Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.
Citation Information
Patent Citations
Method and device for identifying Internet assets based on fingerprints, and medium
CN116015745A
Network penetration result verification method, system and device based on data implicit transmission algorithm and storage medium
CN116248313A
Real-time analytical queries of a document store
US20190361897A1