CMS identification method based on limited request packet detection and multi-feature detection

By constructing multiple scan content regions and feature regions based on limited request packet detection and multi-feature detection methods, the existing CMS recognition methods reduce scanning speed and inaccurate detection results when the feature library is enlarged, and more accurate and efficient CMS recognition is achieved.

CN119996031AActive Publication Date: 2025-05-13JILIN PROVINCE JILIN XIANGYUN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510240054.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing CMS recognition methods lead to a decrease in scanning speed and inaccurate detection results when the feature library is enlarged, and the feature matching method has a high risk of underreport.

Method used

CMS recognition method based on limited request packet detection and multi-feature detection is adopted. By obtaining home page analysis, crawling content and resource files, calculating MD5 and MinHash, building multiple scan content regions and feature regions, and scanning matching recognition is performed within a limited range.

Benefits of technology

It effectively controls the scanning speed and the number of packets sent, reduces the risk of inaccurate detection results, improves the accuracy and detection rate of detection results, and reduces the probability of missed reports.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a CMS (Content Management System) identification method based on limited request packet detection and multi-feature detection, and belongs to the technical field of network security. The problems that in the CMS recognition process, the scanning speed is reduced along with increasing of a feature library, the detection result is inaccurate, vulnerabilities exist and the like are solved. The CMS recognition method based on limited request packet detection and multi-feature detection comprises the steps that home page content of a detection target is obtained through an http request sent to a home page url; the method comprises the following steps: crawling home page content, constructing a plurality of scanning content regions and feature regions, and constructing a complete scanning context; and scanning, matching and identifying the context by using the added CMS fingerprint database recording file in a rule recording manner, and generating an accurate scanning and identifying result. According to the method, the scanning range is standardized, so that the scanning is performed in a limited range, and the scanning speed is not reduced along with the increase of the feature library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a CMS identification method based on limited request packet detection and multi-feature detection, which belongs to the technical field of network security. Background Art

[0002] Content Management System, or CMS for short. A content management system is a programming language that runs on the server to manage and maintain the columns, content, and templates of a website. System website CMS usually includes multiple CMS programs that implement different functions. Currently, CMS identification mainly relies on the fingerprint library, which stores a large number of CMS fingerprints. By extracting the CMS fingerprint of the web page to be identified and comparing the extracted CMS fingerprint with the fingerprint in the fingerprint library, the CMS category of the web page can be determined.

[0003] In the current network security industry, most methods for identifying CMS use single-feature or limited-feature detection technology, such as homepage source code, homepage header, static file content, static file MD5 and other features, and feature matching uses regular or identical expressions for detection. However, the existing detection methods still have the following problems:

[0004] 1. With the increase of feature records in the feature library, the number of detection data packets sent by the detection engine will also increase. When the feature value is too large, the scanning and recognition speed will be reduced. On the other hand, excessive packet sending will cause excessive pressure on the detected website and affect the normal operation of the business. At the same time, if the detected website is protected by security equipment, the probability of being intercepted by the security equipment will increase, resulting in inaccurate detection results.

[0005] 2. The feature equality or regular matching method has a high rate of false negatives and cannot effectively detect CMS systems that have been redeveloped or modified without features. Summary of the invention

[0006] In view of the shortcomings of the existing technology, in order to avoid the problems of reduced scanning speed, inaccurate detection results, and high vulnerability in feature matching as the feature library increases during the CMS identification process, a CMS identification method based on limited request packet detection and multi-feature detection is proposed.

[0007] The CMS identification method based on limited request packet detection and multi-feature detection is characterized by comprising the following steps:

[0008] S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page;

[0009] S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions;

[0010] S3, build a scan context, combining multiple scan content regions and feature regions in S2 to build a complete scan context;

[0011] S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.

[0012] Furthermore, in the step S1, the homepage content is the original data of the homepage obtained, and these data are usually in HTML format, including various elements and reference files required to construct the homepage;

[0013] Parse the acquired homepage content, identify and extract various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and picture tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files;

[0014] The image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.

[0015] Furthermore, in the step S2, the parsed HTML content is crawled, and multiple content regions and feature regions are constructed through content acquisition and feature calculation, which are specifically as follows:

[0016] (1) Extract HTML tags to form label regions;

[0017] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in HTML to form a filelist region;

[0018] (3) Download all files in the filelist region to form the file region;

[0019] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;

[0020] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;

[0021] (6) Extract the external link URL in HTML to form a link region;

[0022] Furthermore, in the step S3, a scanning context is constructed, which specifically includes: creating an initial scanning context framework, incorporating the multiple scanning content regions and feature regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple scanning content regions and feature regions incorporated, and then constructing a scanning context associated with the step S2.

[0023] Furthermore, in step S4, the rule recording method adopts a points system, and the match is successful if the specified points are met; wherein the rule record specifies the points required for the hit, which are specified by the @Match field in the record, and the @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained, and if there is no hit, no points are obtained. The points of all hit detection rules are added together, and if the total points are greater than or equal to the points of @Match4, the hit is recorded and an accurate scanning and recognition result is generated;

[0024] Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.

[0025] The detection rules specified by @Pattern include:

[0026] @Pattern 2region:label;

[0027] @Pattern 1region:file;

[0028] @Pattern 1region:filelist;

[0029] @Pattern 1region:md5;

[0030] @Pattern 1region:minihash;

[0031] @Pattern 1region:link;

[0032] Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.

[0033] The beneficial effects of the present invention are as follows: 1. First, by constructing multiple scanning content regions and feature regions through the parsed HTML content, the multiple scanning content regions and feature regions are then incorporated into the scanning context framework to establish their related detection environment, and to standardize the scanning range, thereby stipulating that the scanning is performed within a limited range. The possibility of reducing the scanning speed due to an increase in the number of acquired scanning elements will not occur as the scanning records increase. On the other hand, the number of data packets sent when scanning the scanning target is effectively controlled, thereby reducing the impact on the target, reducing the risk of inaccurate detection results, and effectively bypassing the target's protective measures.

[0034] 2. Construct multiple scanning content regions and feature regions, clarify the scanning scope, form a multi-feature and multi-region scanning matching method, thereby reducing the probability of missed reports and making the detection results more accurate. Compared with the existing technology, the single type of feature compared to the traditional scanning method improves the flexibility of scanning records and can further improve the detection rate. DETAILED DESCRIPTION

[0035] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.

[0036] This embodiment provides a technical solution: a CMS identification method based on limited request packet detection and multi-feature detection includes the following steps:

[0037] S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page;

[0038] The home page content is the original data of the home page, which is usually in HTML format and contains various elements and reference files required to build the home page.

[0039] Parse the obtained homepage content, identify and extract the various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and image tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files;

[0040] Image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.

[0041] Specifically, for https: / / www.example.comMonitoring detailed steps: send the scanning device to the desired scan through the URL http: / / www.example.com Request and obtain the home page content, usually in HTML format;

[0042] S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions;

[0043] Crawl the parsed HTML content and construct multiple scanning content regions and feature regions. Construct multiple content regions and feature regions through content acquisition and feature calculation. The details are as follows:

[0044] (1) Extract HTML tags to form label regions;

[0045] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in HTML to form a filelist region;

[0046] (3) Download all files in the filelist region to form the file region;

[0047] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;

[0048] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;

[0049] (6) Extract the external link URL in HTML to form a link region;

[0050] Specifically, the HTML content is parsed and multiple scanning content regions and feature regions are constructed, including:

[0051] (1) Extract HTML tags to form label regions;

[0052] (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in html to form filelistregion;

[0053] (3) Download all files in the filelist region to form the file region;

[0054] (4) Calculate the md5 of the file content in the filelist region to form the md5 region;

[0055] (5) Calculate the minihash of the file content in the filelist region to form a minihash region;

[0056] (6) Extract the external link URLs in HTML to form a link region.

[0057] By constructing multiple scanning content regions and feature region combinations through a limited number of http detection packets, there is no need to send additional request packets during the scanning and matching feature process, which standardizes the scanning range during the CMS fingerprint recognition process and further stipulates that the scanning is performed within a limited range.

[0058] S3, build a scan context, combining multiple scan content regions and feature regions in S2 to build a complete scan context;

[0059] Constructing a scanning context specifically includes: creating an initial scanning context framework, incorporating the multiple scanning content regions and feature regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple scan regions included, and then constructing a scanning context associated with the S2 step.

[0060] Specifically, an initial scanning context framework is created, and the multiple scanning content regions and feature regions constructed in the above step S2, including label region, filelist region, file region, md5 region, minihash region and link region, are incorporated into the initial scanning context framework, so as to clarify that the scanning operation is specifically performed on the multiple scanning content regions and feature regions incorporated, and then a scanning context associated with step S2 is constructed, so that the scanning speed will not decrease with the increase of feature records in the feature library, thereby avoiding inaccurate detection results and the problem of underreporting.

[0061] S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.

[0062] The rule recording method adopts a points system. If the specified points are met, the match is successful. The rule record specifies the points required for the hit, which is specified by the @Match field in the record. The @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained. If there is no hit, no points are obtained. The points of all hit detection rules are added together. If the total points are greater than or equal to the points of @Match 4, the hit is recorded and an accurate scanning and recognition result is generated.

[0063] Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.

[0064] The detection rules specified by @Pattern include:

[0065] @Pattern 2region:label;

[0066] @Pattern 1region:file;

[0067] @Pattern 1region:filelist;

[0068] @Pattern 1region:md5;

[0069] @Pattern 1region:minihash;

[0070] @Pattern 1region:link;

[0071] Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.

[0072] The score required for a specified hit in the record is specified by the @Match field in the record. The @Pattern in the record specifies a detection rule. If the rule is hit, the specified score is obtained. If there is no hit, no score is obtained. The scores of all hit rules are added together. If the total score is greater than or equal to the score in @Match, the record is hit and an accurate scanning and recognition result is generated.

[0073] Specifically, the record example of the embodiment is:

[0074] @Match 4

[0075] @Pattern 2" "!region:lable;

[0076] @Pattern 1 "xxxs"! path"xxx.js"! region:file;

[0077] @Pattern 1"ruoyi.css"! region:filelist;

[0078] @Pattern 1"2221f431e845d608b9bb71c07c7a6a18"! path". / js / index.js"! region:md5;

[0079] @Pattern 1"11111"! path". / js / index.js"! Similarity 0.8! region:minihash;

[0080] @Pattern 1"http: / / xxx.com"! region:link;

[0081] Use the added CMS fingerprint library record file to scan and compare with the context constructed above. In the record, @Pattern2" "!region:label

[0082] Matches " in the html label tag "If you hit it, you get 2 points, if you miss it, you get no points;

[0083] @Pattern 1 "xxxs"! path"xxx.js"! region:file

[0084] Match "xxxs" in the xxx.js content in the file region and get 1 point;

[0085] @Pattern 1"ruoyi.css"! region:filelist

[0086] Check whether the filelist region contains the ruoyi.css file. If it does, you will get 1 point.

[0087] @Pattern 1"2221f431e845d608b9bb71c07c7a6a18"! path". / js / index.js"! region:md5

[0088] If the md5 of the file ". / js / index.js" is matched in the md5 region

[0089] "2221f431e845d608b9bb71c07c7a6a18" will get 1 point if hit;

[0090] @Pattern 1"11111"! path". / js / index.js"! Similarity 0.8! region:minihash

[0091] Match ". / js / index.js" in the minhash region to see if the similarity between the minihash value and "11111" is greater than or equal to 0.8. If so, a point is obtained.

[0092] @Pattern 1"http: / / xxx.com"! region:link

[0093] In the link region, check whether the external link "http: / / xxx.com" is included. If it is matched, a point is obtained. The points of all the matches are added together. If it is greater than or equal to the score specified in @Match, the rule is matched, and an accurate scanning and recognition result is output.

[0094] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A CMS identification method based on limited request packet detection and multi-feature detection, characterized in that: The following steps are involved: S1. Get the home page analysis, obtain the home page content of the detection target by sending an http request to the home page URL, and parse out all the elements and referenced files contained in the home page; S2. Crawl all the content and resource files parsed from the homepage content, download and save them in sequence, calculate the MD5 and MinHash of the files, and construct multiple scanning content regions and feature regions; S3, build a scan context, combining multiple scan regions and feature regions in S2 to build a complete scan context; S4. Use the added CMS fingerprint library record file to scan and match the context through rule recording, and generate accurate scanning and recognition results.

2. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S1, the home page content is the original data of the home page obtained, and this data is usually in HTML format, including various elements and reference files required to build the home page; Parse the acquired homepage content, identify and extract various elements and reference files contained therein, wherein the various elements include but are not limited to icons (favicon), titles, paragraphs, links and picture tags in the webpage; the reference files include but are not limited to resource files, JavaScript (.js) files, cascading style sheets (.css) files and image files; The image files include but are not limited to files with suffixes of .jpg, .jpeg, .png, .ico, and .gif.

3. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S2, the parsed HTML content is crawled, and multiple content regions and feature regions are constructed through content acquisition and feature calculation, which are specifically as follows: (1) Extract HTML tags to form label regions; (2) Extract the js, css, icon, jpg, jpeg, png, ico, and gif file paths contained in html to form filelistregion; (3) Download all files in the filelist region to form the file region; (4) Calculate the md5 of the file content in the filelist region to form the md5 region; (5) Calculate the minihash of the file content in the filelist region to form a minihash region; (6) Extract the external link URLs in HTML to form a link region.

4. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In the step S3, a scanning context is constructed, which specifically includes: creating an initial scanning context framework, incorporating the multiple scanning regions constructed in the above step S2 into the scanning context, and then clarifying that the scanning operation is specifically performed on the multiple included scanning regions, and on the basis of the included scanning regions, integrating the calculated features into the scanning context, and then constructing a scanning context associated with the step S2.

5. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 1 is characterized in that: In step S4, the rule recording method adopts a points system, and the match is successful if the specified points are met; the rule record specifies the points required for the hit, which is specified by the @Match field in the record, and the @Pattern in the record specifies a detection rule. If a detection rule is hit, the specified points are obtained, and if there is no hit, no points are obtained. The points of all hit detection rules are added together. If the total points are greater than or equal to the points of @Match 4, the hit is recorded and an accurate scanning and recognition result is generated; Among them, @Match 4 specifies that the sum of the points in the hit rule must be greater than or equal to 4.

6. The CMS identification method based on limited request packet detection and multi-feature detection according to claim 5 is characterized in that: The detection rules specified by @Pattern include: @Pattern 2region:label; @Pattern 1region:file; @Pattern 1region:filelist; @Pattern 1region:md5; @Pattern 1region:minihash; @Pattern 1region:link; Among them, @Pattern 2 means that 2 points can be obtained if the content in this detection rule is hit, and @Pattern 1 means that 1 point can be obtained if the content in this detection rule is hit.

Citation Information

Patent Citations

  • Method and device for identifying Internet assets based on fingerprints, and medium

    CN116015745A

  • Network penetration result verification method, system and device based on data implicit transmission algorithm and storage medium

    CN116248313A

  • Real-time analytical queries of a document store

    US20190361897A1