Web crawler protection method based on device fingerprint and dynamic protection
By determining the characteristic status based on the page link distribution value of the target website and the relevant page difference value, and selecting targeted crawler analysis protection methods, the problem of poor crawler protection efficiency in the existing technology is solved, and more efficient and accurate protection effects are achieved.
Patent Information
- Application Number
- CN202510022459.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The prior art cannot select targeted crawler protection strategies based on actual page conditions, resulting in poor crawler protection efficiency.
By determining the characteristic status of the target website based on the page link distribution value of the target website and the relevant page difference value, different crawler analysis and protection methods are selected, including analysis and protection for user operation fingerprints or dynamic protection for target websites.
The efficiency of crawler protection has been improved, making the choice of protection methods more in line with the actual work scenario, and the accuracy and effectiveness have been improved.
Smart Images

Figure CN120046146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crawler protection, and in particular to a Web crawler protection method based on device fingerprint and dynamic protection. Background Art
[0002] A Web crawler is an automated program or script that usually gradually crawls the links in a web page and recursively accesses these links until a certain condition is met. To cope with the attacks of malicious crawlers, many websites will take a series of crawler protection measures, such as using device fingerprint monitoring, verification codes, and restricting access frequencies. These measures are designed to protect the security and stability of the website and prevent malicious crawlers from over-crawling and damaging the website. However, the current crawler protection measures are poor. For example, in current device fingerprint monitoring, it is possible to determine whether there is a crawler risk based on the user's mouse trajectory and operation rules. However, it is impossible to formulate different device fingerprint monitoring strategies according to the actual situation of the Web page, resulting in poor protection effects.
[0003] Patent publication number CN109189660A discloses a crawler recognition method based on user mouse interaction behavior, including the following steps: Step (1), sampling data of user behavior; Step (2), thinning and encrypting the sampling data in Step (1); Step (3), performing trajectory analysis; Step (4), after discovering that the user is a malicious user, performing a blocking process; among them, in the trajectory analysis step, multi-dimensional analysis is performed, including whether the user's mouse movement trajectory is smooth and continuous, whether the scrolling distance of the page per unit time is reasonable, and a threshold and a weight value are set for each dimension. When the sampling information of a single dimension breaks through the threshold, it is determined as a malicious user, or when the result value after weight calculation of multiple dimensions is lower than a certain threshold, it is determined as a malicious user; thus, it can be seen that the comparative document has the following problems: It does not consider that the link distribution states corresponding to different pages are different. For example, for pages with different link distribution densities, the actual smoothness and continuity of the user's mouse movement trajectory and the scrolling distance of the page per unit time are different, which is prone to misjudgment or wrong judgment, and it is impossible to make active crawler protection. Summary of the Invention
[0004] Therefore, the present invention provides a Web crawler protection method based on device fingerprint and dynamic protection to overcome the problem in the prior art that the crawler protection efficiency is poor because different targeted crawler protection strategies cannot be selected according to the actual page situation.
[0005] To achieve the above object, the present invention provides a Web crawler protection method based on device fingerprint and dynamic protection, including:
[0006] Determining the characteristic state of the target website according to the page link distribution value of the target website and the relevant page difference value;
[0007] Determine the crawler analysis protection method based on the feature status of the target website as analyzing and protecting the user operation fingerprint or dynamically protecting the target website;
[0008] In the analysis and protection of the user operation fingerprint, determine whether to issue a crawler protection warning according to the comparison result of the link information difference degree and the preset link information difference degree, or according to the information-related reference value and the search depth reference value, or according to the mouse track anomaly reference value and the mouse sliding anomaly reference value;
[0009] In the dynamic protection of the target website, determine the dynamic protection requirement status according to the number of sensitive access pages and the resource utilization rate, and determine the dynamic protection strategy according to the dynamic protection requirement status.
[0010] Further, when the feature status of the target website is that the page link distribution value is within the first preset distribution value range or the relevant page difference value is within the first preset difference value range, analyze and protect the user operation fingerprint, including:
[0011] Detect the link information difference degree;
[0012] If the link information difference degree is greater than the preset link information difference degree, determine whether to issue a crawler protection warning according to the information-related reference value and the search depth reference value;
[0013] If the link information difference degree is less than or equal to the preset link information difference degree, determine whether to issue a crawler protection warning according to the mouse track anomaly reference value and the mouse sliding anomaly reference value.
[0014] Further, the confirmation method of the link information difference degree includes:
[0015] Detect the validity of the link name of the target website;
[0016] If the link name validity is greater than the preset link name validity, the link information difference degree is determined according to the number of similar main page link names;
[0017] If the link name validity is less than or equal to the preset link name validity, the link information difference degree is determined according to the number of similar link names of the relevant pages.
[0018] Further, the confirmation method of the mouse track anomaly reference value includes:
[0019] Detect the distribution similarity corresponding to the target website and the historical record, and the distribution similarity is determined according to the link distance and the number of links;
[0020] If the distribution similarity is greater than the preset distribution similarity, determine the abnormal reference value of the mouse trajectory according to the mouse trajectory similarity;
[0021] If the distribution similarity is less than or equal to the preset distribution similarity, determine the abnormal reference value of the mouse trajectory according to the number of mouse disconnections.
[0022] Further, the confirmation method of the abnormal reference value of mouse sliding includes:
[0023] Obtain a number of usage frames according to the preset selection rules;
[0024] Detect the number of images in each usage frame;
[0025] If the number of images is within the first preset effective image number range, determine that the abnormal reference value of mouse sliding is the preset response reference value;
[0026] If the number of images is within the second preset effective image number range, determine the abnormal reference value of mouse sliding according to the image repetition value.
[0027] Further, when the number of images is within the second preset effective image number range, the abnormal reference value of mouse sliding has a positive correlation with the image repetition value.
[0028] Further, the information-related reference value is determined based on the search deviation degree;
[0029] The search deviation degree has a negative correlation with the number of keyword combinations on the user-requested page.
[0030] Further, the search depth reference value is determined based on the number of vertically searched pages;
[0031] The search depth reference value has a positive correlation with the number of vertically searched pages.
[0032] Further, when the characteristic state of the target website is that the page link distribution value is within the second preset distribution value range and the relevant page difference value is within the second preset difference value range, perform dynamic protection on the target website, including:
[0033] Detect the number of sensitive access pages and the resource utilization rate;
[0034] Determine the dynamic protection requirement status according to the number of sensitive access pages and the resource utilization rate;
[0035] Determine the dynamic protection strategy according to the dynamic protection requirement status;
[0036] If the dynamic protection requirement status is that the number of sensitive access pages is within the first preset sensitive access page number range or the resource utilization rate is within the first resource utilization rate range, the dynamic protection strategy is to perform port hopping at a fixed frequency;
[0037] If the dynamic protection requirement status is that the number of sensitive access pages is within the second preset range of the number of sensitive access pages and the resource utilization rate is within the second resource utilization rate range, the dynamic protection policy is to perform port hopping at an increasing hopping frequency.
[0038] Furthermore, the increasing hopping frequency has a positive correlation with the number of sensitive access pages;
[0039] The increasing hopping frequency has a negative correlation with the resource utilization rate.
[0040] Compared with the prior art, the beneficial effect of the present invention is that in the technical solution of the present invention, the characteristic state of the target website is determined according to the page link distribution value and the related page difference value of the target website, and the distribution of the links of the pages in the target website is reflected through the characteristic state, and different crawler analysis protection methods are correspondingly selected, so that the selection of the crawler analysis protection method is more in line with the actual working scenario, thereby improving the crawler protection efficiency of the present invention.
[0041] Furthermore, in the analysis and protection of the user operation fingerprint in the present invention, according to the comparison result between the link information difference degree and the preset link information difference degree, it is selected to determine whether to perform crawler protection warning according to the information correlation reference value and the search depth reference value or to determine whether to perform crawler protection warning according to the mouse track anomaly reference value and the mouse sliding anomaly reference value. The content difference degree between the corresponding contents of the links corresponding to the target website is reflected through the link information difference degree, so that the selection of the criterion for determining whether to perform crawler protection warning is more accurate.
[0042] Furthermore, when the link information difference degree in the present invention is greater than the preset link information difference degree, it is determined whether to perform crawler protection warning according to the information correlation reference value and the search depth reference value. By taking advantage of the drawback that most crawlers cannot recognize the relevance between pages, the relevance degree of each page during the user page request process is reflected through the information correlation reference value, thereby reflecting the risk of crawler attack, and thus improving the crawler recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the Web crawler protection method based on device fingerprint and dynamic protection of the present invention;
[0044] Figure 2 It is a schematic diagram of determining the dynamic protection requirement status according to the number of sensitive access pages and the resource utilization rate of the present invention;
[0045] Figure 3 It is a flowchart of analyzing and protecting the user operation fingerprint of the present invention;
[0046] Figure 4This is a flowchart of the method for confirming the difference degree of link information of the present invention. Detailed implementation manners
[0047] In order to make the objectives and advantages of the present invention more clear and understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The preferred implementation manners of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0049] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0050] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0051] Please refer to Figures 1 to 4 As shown, the present invention provides a Web crawler protection method based on device fingerprint and dynamic protection, including:
[0052] Determine the characteristic state of the target website according to the page link distribution value and the relevant page difference value of the target website;
[0053] Based on the characteristic state of the target website, determine that the crawler analysis protection method is to analyze and protect the user operation fingerprint or to perform dynamic protection on the target website;
[0054] In the analysis and protection of the user operation fingerprint, determine whether to perform crawler protection warning according to the comparison result of the link information difference degree and the preset link information difference degree, according to the information correlation reference value and the search depth reference value, or according to the mouse track anomaly reference value and the mouse sliding anomaly reference value;
[0055] In the dynamic protection of a target website, the dynamic protection requirement status is determined according to the number of sensitive access pages and the resource utilization rate, and the dynamic protection strategy is determined according to the dynamic protection requirement status.
[0056] In the present invention, the target website is a website that needs to be protected against crawlers. The main page of the target website is the page currently in use, and the related pages of the target website are the pages searched through the links of the main page of the target website or the links of the related pages. The single page described in the present invention can be a single main page or a related page. The method for confirming the link distribution value corresponding to a single page is to calculate the reference distance corresponding to each link of the page, and record the average value of the reference distances as the page link distribution value. The link distribution value of the target website is the link distribution value corresponding to the main page of the target website. For a single link, the corresponding reference distance is the minimum value among the distances between this link and other links. The method for confirming the related page difference value is to detect the link structure characteristic values of each related page of the target website, calculate the absolute value of the difference between each link structure characteristic value and the reference average value respectively, and record this absolute value as the link structure characteristic difference. The number of link structure characteristic differences greater than the preset link structure characteristic difference is recorded as the related page difference value. The reference average value is the average value of each link structure characteristic value. For a single page, its corresponding link structure characteristic value = the number of links corresponding to the page × α1 + the proportion of the picture area of the page × α2 + the page link distribution value of the page × α3, where α1 is the first calculation coefficient, α2 is the second calculation coefficient, and α3 is the third calculation coefficient. It can be understood that the values of α1, α2, and α3 can be determined according to the user's attention degree to the number of links corresponding to the page, the proportion of the picture area, or the link structure characteristic value. For example, if the user's past experience determines that the number of links corresponding to the page has a greater impact on the page similarity degree, the value of the first calculation coefficient is larger. Provide a set of values: α1 = 0.4, α2 = 0.3, α3 = 0.3. The preset link structure characteristic difference is 30% of the reference average value. It can be understood that the greater the user's determination accuracy for the page similarity degree, the larger the preset link structure characteristic difference. The present invention reflects the page similarity degree and the link distribution density of the target website through the page link distribution value and the related page difference value, and then determines the corresponding protection method, avoiding the problem in the prior art that for pages with a large page similarity degree and a large link distribution density, it is impossible to effectively protect against crawlers according to the user device operation fingerprint.
[0057] Specifically, when the characteristic status of the target website is that the page link distribution value is within the first preset distribution value range or the related page difference value is within the first preset difference value range, analyze and protect the user operation fingerprint, including:
[0058] Detect the difference degree of link information;
[0059] If the difference degree of link information is greater than the preset difference degree of link information, determine whether to give a crawler protection warning according to the information correlation reference value and the search depth reference value;
[0060] If the difference degree of link information is less than or equal to the preset difference degree of link information, determine whether to give a crawler protection warning according to the mouse track anomaly reference value and the mouse sliding anomaly reference value;
[0061] The page link distribution value being within the first preset distribution value range or the relevant page difference value being within the first preset difference value range is recorded as the first characteristic state
[0062] Specifically, the values within the first preset distribution value range of the present invention are all greater than the preset distribution value, the values within the second preset distribution value range are all less than or equal to the preset distribution value, the values within the first preset difference value range are all less than the preset difference value, the values within the second preset difference value range are all greater than or equal to the preset difference value. For the values of the preset distribution value and the preset difference value, the user can set them according to the actual application scenario. It is possible to perform data cleaning on the page link distribution value and the relevant page difference degree corresponding to the historical crawler protection records that meet the user's needs to remove the outliers. The method for confirming the outliers can be the Z-Score method or the IQR method. Denote the average value of the page link distribution value after removing the outliers as the preset distribution value, and denote the average value of the relevant page difference degree after removing the outliers as the preset difference degree; The present invention applies historical crawler protection records. Any one of the historical records in the historical crawler protection records records at least one crawler protection usage process, and each historical crawler protection record corresponds to a qualified mark. The qualified mark shows whether the crawler protection effect meets the user's needs. The qualified mark can be manually recorded. Among them, it can be understood that the user can determine whether the crawler protection effect meets the requirements according to the accuracy rate of the crawler protection by himself / herself, which will not be elaborated here.
[0063] Specifically, the confirmation methods of the difference degree of link information include:
[0064] Detect the validity of the link name of the target website;
[0065] If the link name validity is greater than the preset link name validity, the difference degree of link information is determined according to the number of similar link names on the main page;
[0066] If the link name validity is less than or equal to the preset link name validity, the difference degree of link information is determined according to the number of similar link names on the relevant pages.
[0067] Specifically, the number of recognizable keywords in the link name of the main page of the target website is recorded as the link name validity. All link names of the main page of the target website are extracted, and the keywords in the link name are recognized through a deep learning network and recorded as recognizable keywords. The value range of the link name validity is preset and can be set by the user according to actual needs. Data cleaning is performed on the link name validity corresponding to the historical crawler protection records that meet the user's needs to remove outliers. The method for confirming outliers can be the Z-Score method or the IQR method. The average value of the link name validity after removing outliers is recorded as the preset link name validity. It can be understood that the link name validity can reflect whether the link name of the main page of the target website can meet the calculation requirements of the link information difference degree. It can be understood that the smaller the link name validity, the smaller the accuracy of the link information difference degree determined according to the similarity number of the link names of the main page.
[0068] If there are the same keywords in both link names, then the same keywords are recorded as duplicate keywords, and the number of duplicate keywords is the similarity number.
[0069] The confirmation method of the link information difference degree is as follows: when determining according to the similarity number of the link names of the main page, (the number of recognizable keywords in the link name of the main page - the similarity number of the link names of the main page) / the number of recognizable keywords in the link name of the main page is recorded as the link information difference degree. When the link information difference degree is determined according to the similarity number of the link names of the relevant pages, (the number of recognizable keywords in the link name of the relevant page - the similarity number of the link names of the relevant page) / the number of recognizable keywords in the link name of the relevant page is recorded as the link information difference degree. The value range of the preset link information difference degree is set by the user according to actual crawler protection needs and past experience. The greater the user's demand for crawler protection performance, the greater the value of the preset link information difference degree. To avoid misjudgment in determining whether to issue a crawler protection warning based on the information correlation reference value and the search depth reference value, a value of the preset link information difference degree is provided. The preset link information difference degree is 70% of the number of recognizable keywords in the link name of the main page or the relevant page. Among them, in the confirmation of the preset link information difference degree, the selection of the main page and the relevant page needs to be determined according to the calculation method of the link information difference degree, which is easy to understand for those skilled in the art and will not be elaborated here.
[0070] Specifically, the method for confirming the mouse trajectory anomaly reference value includes:
[0071] Detect the distribution similarity corresponding to the target website and the historical records, and the distribution similarity is determined according to the link distance and the number of links.
[0072] If the distribution similarity is greater than the preset distribution similarity, determine the mouse track anomaly reference value according to the mouse track similarity;
[0073] If the distribution similarity is less than or equal to the preset distribution similarity, determine the mouse track anomaly reference value according to the number of mouse disconnections.
[0074] Specifically, the way to confirm the distribution similarity is to divide the main page of the target website into several sub-regions with the same area. For the main page and any reference page, if the number of similar sub-regions between the two pages is greater than the preset similar number, then the reference page is the similar page of the main page, and record the number of similar pages of the main page as the distribution similarity. For a sub-region and a reference historical sub-region, if the absolute value of the difference in the number of links corresponding to the two regions is less than 3 and the absolute value of the difference in the maximum link distance is less than 10% of the maximum link distance, then the sub-region and the reference historical sub-region are similar sub-regions. The preset similar number = the number of sub-regions of the main page × ζ. The greater the user's requirement for the determination accuracy of the similarity degree of the page, the greater the value of ζ. However, it should be noted that ζ should be greater than 1 and less than 2. Provide a value of ζ, ζ = 1.6. The greater the user's requirement for the accuracy of crawler protection according to historical records, the greater the value of the preset distribution similarity. In the present invention, a historical record includes a page (i.e., the reference page) requested by the user historically, and its related information. The related information includes the link distance and the number of links within the historical sub-region of the reference page, and the mouse operation track. The historical sub-region is several sub-regions with the same area obtained by dividing the reference page of the historical record. The recording of the mouse operation track can be realized by tools with mouse tracking recording functions such as Mouse Tracks.
[0075] Specifically, the way to confirm the mouse sliding anomaly reference value includes:
[0076] Obtain several usage frames according to the preset selection rule;
[0077] Detect the number of images in each usage frame;
[0078] If the number of images is within the first preset valid image number range, determine that the mouse sliding anomaly reference value is the preset response reference value;
[0079] If the number of images is within the second preset valid image number range, determine the mouse sliding anomaly reference value according to the image repetition value.
[0080] Among them, the preset selection rule is to capture an image of the mouse reference range in the page currently used by the user every 5 seconds and record it as a usage frame. The mouse reference range is a circle centered on the current mouse position, and the area of this circle is 15% of the display area of the current page. The area of the mouse reference range can be determined by a deep learning network learning from historical crawler protection records. How to set the training set and validation set, and how to perform model learning and validation are content easily understood by those skilled in the art. The values within the first preset effective image quantity range are all less than the preset effective image quantity, and the values within the second preset effective image quantity range are all greater than or equal to the preset effective image quantity. The number of images in the usage frame is the number of non-touching images other than links. The value of the preset effective image quantity can be set by the user according to the actual application scenario. Data cleaning can be performed on the historical crawler protection records corresponding to the pages that meet the user's needs to remove outliers, and the average value of the image quantity after removing outliers is recorded as the preset effective image quantity. The preset response reference value is 0.
[0081] Specifically, when the number of images is within the second preset effective image quantity range, the mouse sliding anomaly reference value and the image repetition value have a positive correlation.
[0082] The image repetition value is the maximum number of times each image in the collected usage frames appears repeatedly as of the current moment. By utilizing the characteristic that the mouse sliding rule during the crawler usage process cannot be changed according to the image distribution, anomaly recognition is performed, thereby improving the crawler recognition efficiency.
[0083] Specifically, the information correlation reference value is determined based on the search deviation degree;
[0084] The search deviation degree has a negative correlation with the number of keyword combinations on the user-requested page.
[0085] The information correlation reference value = 1 / search deviation degree. The user-requested page is the M pages that the user has most recently requested as of the current moment. The confirmation method for the number of keyword combinations is as follows: Extract the request links corresponding to each user-requested page. Through this request link, the user-requested page can be entered. Extract the recognizable keywords of each request link, and detect the number of times any two recognizable keywords appear simultaneously in the historical healthy search records and record it as the combination number. The average value of the combination numbers is recorded as the number of keyword combinations. The value of M can be set by the user according to the actual application scenario and will not be elaborated here. M is preferably 10.
[0086] The historical crawler protection records without detecting crawlers are recorded as historical healthy search records.
[0087] Specifically, the search depth reference value is determined based on the number of vertically searched pages;
[0088] The search depth reference value has a positive correlation with the number of vertical search pages.
[0089] The method for confirming the number of vertical search pages is to detect the main page currently used by the user, detect the number of indirect links from the initial page of the target website to the main page, and record this number of indirect links as the number of vertical search pages. The number of indirect links is the number of links clicked from the initial page of the target website to the main page.
[0090] Specifically, when the characteristic state of the target website is that the page link distribution value is within the range of the second preset distribution value and the relevant page difference value is within the range of the second preset difference value, dynamic protection is performed on the target website, including:
[0091] Detect the number of sensitive access pages and the resource utilization rate;
[0092] Determine the dynamic protection requirement status according to the number of sensitive access pages and the resource utilization rate;
[0093] Determine the dynamic protection strategy according to the dynamic protection requirement status;
[0094] If the dynamic protection requirement status is that the number of sensitive access pages is within the range of the first preset number of sensitive access pages or the resource utilization rate is within the range of the first resource utilization rate, the dynamic protection strategy is to perform port hopping at a fixed frequency;
[0095] If the dynamic protection requirement status is that the number of sensitive access pages is within the range of the second preset number of sensitive access pages and the resource utilization rate is within the range of the second resource utilization rate, the dynamic protection strategy is to perform port hopping at an increasing hopping frequency.
[0096] Specifically, whether a page is a sensitive page is set by the user himself. It can be understood that the user can determine whether a page is set as a sensitive page according to factors including but not limited to information value, information quantity, and the number of links in the page, which will not be elaborated here. The number of sensitive access pages is the number of sensitive pages requested to be accessed by the user in the most recent monitoring period. Resource utilization rate = (CPU utilization rate / preset CPU utilization rate) + (memory utilization rate / preset memory utilization rate). The values of the preset CPU utilization reference value and the preset memory utilization rate can be set by the user himself. It can be understood that the greater the resource utilization rate, the greater the impact on the Web system service capacity. A value-taking method is provided. The user can record the average value of the CPU utilization reference value and the average value of the memory utilization rate corresponding to the records that meet the user's requirements in the historical crawler protection records as the preset CPU utilization reference value and the preset memory utilization rate respectively. The values within the first resource utilization rate range are all smaller than the preset resource utilization rate, and the values within the second resource utilization rate range are all greater than or equal to the preset resource utilization rate. The user can set the value of the preset resource utilization rate according to actual needs. A value of the preset resource utilization rate is provided. The preset resource utilization rate is preferably 1.2. The values within the first preset number of sensitive access pages range are all smaller than the preset number of sensitive access pages, and the values within the second preset number of sensitive access pages range are all greater than or equal to the preset number of sensitive access pages. The preferred value of the preset number of sensitive access pages is 20. The greater the user's demand for crawler protection, the smaller the value of the preset number of sensitive access pages.
[0097] In the present invention, the Web service system corresponding to the target website is continuously controlled to perform periodic port hopping. The value of the fixed frequency is set by the user according to the crawler protection requirement. The greater the user's crawler protection requirement, the greater the fixed frequency. A value of the fixed frequency is provided. The fixed frequency is 1 time / 30s. When the port hopping is performed at the fixed frequency, the Web service system reallocates the page resources every 30s and assigns a random or pseudo-random port number.
[0098] Specifically, the growth hopping frequency and the number of sensitive access pages have a positive correlation.
[0099] The growth hopping frequency and the resource utilization rate have a negative correlation.
[0100] When the page link distribution value is within the second preset distribution value range and the relevant page difference value is within the second preset difference value range, it is recorded as the second characteristic state.
[0101] It should be noted that the minimum value of the growth hopping frequency should be greater than the fixed frequency. When the port hopping is performed at the growth hopping frequency, the Web service system allocates the page resources according to the growth hopping frequency.
[0102] When determining the abnormal reference value of the mouse trajectory based on the similarity of the mouse trajectories, the abnormal reference value of the mouse trajectory is the similarity of the mouse trajectories, and the similarity of the mouse trajectories = the number of similar trajectories / the total number of trajectories. For two mouse trajectories, if the Euclidean distance reference value corresponding to the two mouse trajectories is less than the preset Euclidean distance reference value, the two mouse trajectories are similar trajectories. In the present invention, 50 mouse trajectories during the user's most recent operation process and the mouse trajectories in the most recent 50 historical records are extracted. A single mouse trajectory is the trajectory of the mouse operation between two link requests. The number of similar trajectories among the 100 mouse trajectories is recorded as the number of similar trajectories. For two mouse trajectories, the method for confirming the corresponding Euclidean distance reference value is as follows: Obtain the coordinate data of the two mouse trajectories, and record the two trajectories as trajectory A and trajectory B respectively. Randomly select n coordinate points on each of trajectory A and trajectory B. Note that there are n coordinate points on both trajectories. Place the left endpoints of the two trajectories at the origin. Each coordinate point contains values in two dimensions, x and y. For each corresponding point in trajectory A and trajectory B, calculate the Euclidean distance between them. Among them,
[0103]
[0104] where XAi is the x coordinate corresponding to the i-th coordinate point on trajectory A, XBi is the x coordinate corresponding to the i-th coordinate point on trajectory B, YAi is the y coordinate corresponding to the i-th coordinate point on trajectory A, and YBi is the y coordinate corresponding to the i-th coordinate point on trajectory B. The calculation formula for the Euclidean distance reference value D0 is:
[0105]
[0106] It can be understood that the smaller D0 is, the more similar the two trajectories are. The value of the preset Euclidean distance reference value can be set by the user according to actual needs. For those skilled in the art, based on the data corresponding to the historical crawler protection records and the actual crawler protection requirements, and on the premise that the calculation method of the Euclidean distance reference value has been provided, setting the preset Euclidean distance reference value is an understandable content and will not be elaborated here. A preferred preset Euclidean distance reference value is provided, with a value of (dmax + dmin) / 2 × 0.5, where dmax is the maximum value in di, dmin is the minimum value in di, and i = 1, 2, 3, ……, n.
[0107] When determining the abnormal reference value of the mouse trajectory based on the number of mouse disconnections, the abnormal reference value of the mouse trajectory = the number of mouse disconnections.
[0108] The information-related reference value, the search depth reference value, the mouse track anomaly reference value, and the mouse slide anomaly reference value are each separately set with a preset comparison threshold, and the corresponding preset comparison thresholds are different. Users can determine based on the records in the historical crawler protection records that meet the user's needs. For example, for the preset comparison threshold corresponding to the information-related reference value, the user can record the average value of the information-related reference values corresponding to the records that meet the user's needs in the historical crawler protection records as the preset comparison threshold, and the user can adjust the preset comparison threshold according to the crawler protection accuracy requirements. The greater the user's need for crawler protection, the smaller the preset comparison threshold.
[0109] When determining whether to issue a crawler protection warning based on the information-related reference value and the search depth reference value, if one of the information-related reference value and the search depth reference value is greater than the corresponding preset comparison threshold, a crawler protection warning is issued.
[0110] When determining whether to issue a crawler protection warning based on the mouse track anomaly reference value and the mouse slide anomaly reference value, if one of the mouse track anomaly reference value and the mouse slide anomaly reference value is greater than the corresponding preset comparison threshold, a crawler protection warning is issued.
[0111] The crawler protection warning is for the Web service system to remind the management personnel to conduct user identity verification and verification to determine whether to perform manual blocking. This is content that is easily understood by those skilled in the art and will not be elaborated here.
[0112] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0113] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A web crawler protection method based on device fingerprint and dynamic protection, characterized in that: include: Determine the characteristic status of the target website according to the page link distribution value and the related page difference value of the target website; Determine the crawler analysis protection method based on the characteristic status of the target website, which is to analyze and protect the user operation fingerprint or to dynamically protect the target website; In the analysis and protection of user operation fingerprints, whether to issue a crawler protection warning is determined based on the comparison result of the link information difference degree with the preset link information difference degree, based on the information-related reference value and the search depth reference value or based on the mouse track abnormality reference value and the mouse sliding abnormality reference value; In dynamic protection of the target website, the dynamic protection demand status is determined according to the number of sensitive access pages and resource utilization, and the dynamic protection strategy is determined according to the dynamic protection demand status.
2. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 1, characterized in that: When the characteristic state of the target website is that the page link distribution value is within the first preset distribution value range or the related page difference value is within the first preset difference value range, analysis and protection are performed on the user operation fingerprint, including: Detect the difference of link information; If the link information difference is greater than the preset link information difference, determine whether to issue a crawler protection warning based on the information-related reference value and the search depth reference value; If the link information difference is less than or equal to the preset link information difference, it is determined whether to issue a crawler protection warning according to the mouse track abnormality reference value and the mouse sliding abnormality reference value.
3. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 2, characterized in that: Methods for confirming the difference of link information include: Check the validity of the link name of the target website; If the link name validity is greater than the preset link name validity, the link information difference is determined based on the number of similarities of the link names of the main page; If the link name validity is less than or equal to the preset link name validity, the link information difference is determined according to the number of similarities of the link names of the relevant pages.
4. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 3, characterized in that: The confirmation methods of abnormal mouse trajectory reference values include: Detecting the distribution similarity between the target website and the historical records, wherein the distribution similarity is determined based on the link distance and the number of links; If the distribution similarity is greater than the preset distribution similarity, determining a mouse trajectory abnormality reference value according to the mouse trajectory similarity; If the distribution similarity is less than or equal to the preset distribution similarity, the mouse track abnormality reference value is determined according to the number of mouse disconnection times.
5. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 4, characterized in that: The confirmation methods of abnormal mouse sliding reference value include: Obtaining a number of use frames according to a preset selection rule; Detect the number of images in each used frame; If the number of images is within the first preset valid image number range, determining the mouse sliding abnormality reference value as a preset response reference value; If the number of images is within a second preset valid image number range, it is determined that the mouse sliding abnormality reference value is determined based on the image repetition value.
6. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 5, characterized in that: When the number of images is within the second preset valid image number range, the mouse sliding abnormality reference value is positively correlated with the image repetition value.
7. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 6, characterized in that: The information-related reference value is determined based on the search deviation degree; There is a negative correlation between the search deviation and the number of keyword combinations in the user's requested page.
8. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 7, characterized in that: The search depth reference value is determined based on the number of vertical search pages; The search depth reference value is positively correlated with the number of vertical search pages.
9. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 8, characterized in that: When the characteristic state of the target website is that the page link distribution value is within the second preset distribution value range and the related page difference value is within the second preset difference value range, dynamic protection is performed on the target website, including: Detect the number of sensitive access pages and resource utilization; Determine the dynamic protection demand status based on the number of sensitive access pages and resource utilization; Determine dynamic protection strategy according to dynamic protection demand status; If the dynamic protection requirement state is that the number of sensitive access pages is within the first preset sensitive access page number range or the resource utilization is within the first resource utilization range, the dynamic protection strategy is to perform port hopping at a fixed frequency; If the dynamic protection requirement state is that the number of sensitive access pages is within the second preset sensitive access page number range and the resource utilization is within the second resource utilization range, the dynamic protection strategy is to perform port hopping with an increasing hopping frequency.
10. The method for protecting against web crawlers based on device fingerprint and dynamic protection according to claim 9, characterized in that: The increasing jump frequency is positively correlated with the number of sensitive access pages; The increase jump frequency is negatively correlated with the resource utilization rate.
Citation Information
Patent Citations
A crawler identification method based on user mouse interaction behavior
CN109189660A
Method and device for protecting websites
CN107666471A
Website asynchronous sequence data intelligent acquisition method based on dynamic self-adaption
CN114297462A
Dynamic protection method of intelligent WEB protection system
CN114499926A
Multi-technology fusion intelligent anti-crawler method and system
CN114969678A