Method, electronic device, storage medium and computer program product for anti-grabbing station
By filtering response instructions in red and black trees and two-way linked lists, the problem of malicious crawlers that cannot effectively identify and prevent IP pools in the prior art, and accurate identification and efficient protection of crawlers are achieved.
Patent Information
- Application Number
- CN202211159655.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-22
AI Technical Summary
Existing anti-crawler methods cannot effectively identify and block malicious crawlers using IP pools, resulting in information leakage and increased server operation costs.
By using the address hash value to determine the filtering rules in the red and black tree, combining the domain name hash value to obtain storage bits in the two-way linked list, filter out the response instructions, and achieve accurate identification and blocking of crawlers.
It improves the recognition accuracy of malicious crawlers and the acquisition speed of response instructions, and reduces unnecessary maintenance costs.
Smart Images

Figure CN115514565B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet security technology, and particularly relates to a method for preventing website scraping, an electronic device, a storage medium, and a computer program product. Background Art
[0002] Web crawlers are an important part of search engines and are used to automatically obtain web page content.
[0003] For a content-driven website, it is inevitable to be visited by web crawlers. The crawling frequency of normal search engine crawlers is relatively reasonable and consumes less website resources. However, malicious crawlers have poor web scraping capabilities and often send hundreds of requests concurrently in a loop to scrape websites. These crawlers can cause devastating blows to websites and have extremely strong destructive power.
[0004] For malicious crawlers, the usual method for preventing website scraping is to count and identify crawler IPs (Internet Protocol Addresses). When a crawler IP accesses the target server exceeding the threshold, it will be considered a malicious crawler IP, and then the access request of this crawler IP will be blocked.
[0005] However, if a malicious crawler selects different IPs from the IP pool to access the target server each time, the anti-crawler method of the related technology will not be able to accurately prevent malicious scraping by hackers, easily leak company information, and increase the operating cost of the target server. Summary of the Invention
[0006] The present disclosure provides a method for preventing website scraping, an electronic device, a storage medium, and a computer program product.
[0007] According to one aspect of the present disclosure, there is provided a method for preventing website scraping, including: determining a filtering rule for address information in a red-black tree corresponding to an address hash value according to the address hash value for characterizing the address information in an access request; obtaining a storage bit of the domain name information in a doubly linked list corresponding to the domain name hash value according to the domain name hash value for characterizing the domain name information in the access request; and screening out a response instruction corresponding to the storage bit in the filtering rule.
[0008] The method of an anti-crawler station according to at least one embodiment of the present disclosure determines a filtering rule for address information in a red-black tree corresponding to an address hash value according to the address hash value used to characterize the address information in an access request, including: performing eigenvalue calculation on the address hash value used to characterize the address information in the access request to obtain an address feature corresponding to the address hash value; determining a first element corresponding to the address information according to the address feature, where the first element contains a filtering rule applicable to the address information; and traversing each filtering rule in the red-black tree of the first element based on the address hash value to screen out a filtering rule adapted to the address information.
[0009] The method of an anti-crawler station according to at least one embodiment of the present disclosure obtains a storage location of domain name information in a doubly linked list corresponding to a domain name hash value according to the domain name hash value used to characterize the domain name information in an access request, including: performing eigenvalue calculation on the domain name hash value used to characterize the domain name information in the access request to obtain a domain name feature corresponding to the domain name information; determining a second element corresponding to the domain name feature according to the domain name feature, where the second element has a storage location corresponding to the domain name feature; and traversing each storage location in the doubly linked list of the second element based on the domain name hash value to screen out a storage location corresponding to the domain name feature.
[0010] The method of an anti-crawler station according to at least one embodiment of the present disclosure, before determining a filtering rule for address information in a red-black tree corresponding to an address hash value according to the address hash value used to characterize the address information in an access request, further includes: performing a hash operation on the address information in the access request to obtain an address hash value used to characterize the address information in the access request.
[0011] The method of an anti-crawler station according to at least one embodiment of the present disclosure, before determining a filtering rule for address information in a red-black tree corresponding to an address hash value according to the address hash value used to characterize the address information in an access request, further includes: in response to an update time, triggering multiple timers respectively to delete the expired rules in the red-black trees of each element.
[0012] The method of an anti-crawler station according to at least one embodiment of the present disclosure, in response to an update time, triggering multiple subprocesses respectively to delete the expired rules in the red-black trees of each element, includes: determining multiple target elements matched by each subprocess according to the number of subprocesses and the number of elements; and in response to an update time, triggering each subprocess respectively to delete the expired rules whose expiration time exceeds the target time according to the expiration time of each filtering rule in the red-black tree of the target element.
[0013] The method of the anti-crawling station according to at least one embodiment of the present disclosure, before determining the filtering rule for the address information in the red-black tree corresponding to the address hash value according to the address hash value used to characterize the address information in the access request, includes: constructing a plurality of elements, including: respectively constructing a red-black tree and a doubly linked list in each element; calculating the rule hash value of each filtering rule; obtaining the rule feature of the filtering rule according to the rule hash value; allocating each filtering rule to the corresponding element according to the rule feature; and storing the filtering rules in the same element into the red-black tree in sequence according to the expiration time of the filtering rule.
[0014] The method of the anti-crawling station according to at least one embodiment of the present disclosure, before obtaining the storage position of the domain name information in the doubly linked list corresponding to the domain name hash value according to the domain name hash value used to characterize the domain name information in the access request, includes: performing a hash operation on the domain name information in the access request to obtain the domain name hash value used to characterize the domain name information in the access request.
[0015] The method of the anti-crawling station according to at least one embodiment of the present disclosure, after screening out the response instruction corresponding to the storage position from the filtering rules, includes: in response to the response instruction for executing the penalty measure, blocking the crawling operation of the request subject of the access request; and in response to the response instruction for allowing access, allowing the crawling operation of the request subject.
[0016] According to another aspect of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method of the anti-crawling station according to any one of the embodiments of the present disclosure.
[0017] According to still another aspect of the present disclosure, there is provided a readable storage medium storing a computer program, and the computer program is suitable for being loaded by a processor to execute the method of the anti-crawling station according to any one of the embodiments of the present disclosure.
[0018] According to yet another aspect of the present disclosure, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the method of the anti-crawling station according to any one of the embodiments of the present disclosure is implemented. Description of the Drawings
[0019] The drawings illustrate exemplary embodiments of the present disclosure and are used together with the description to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure, and the drawings are included in this specification and form a part of this specification.
[0020] Figure 1 It is a flowchart of the method of the anti-crawling station according to an exemplary embodiment of the present disclosure.
[0021] Figure 2 Schematic diagram of the application of the anti-grabbing station method according to an exemplary embodiment of the present disclosure.
[0022] Figure 3 Schematic diagram of an array structure according to an exemplary embodiment of the present disclosure.
[0023] Figure 4 Block diagram of the anti-grabbing station device according to an exemplary embodiment of the present disclosure.
[0024] Description of the Reference Numerals
[0025] 1000 Anti-grabbing station device
[0026] 1002 Filter rule determination module
[0027] 1004 Storage bit determination module
[0028] 1006 Response instruction generation module
[0029] 1100 Bus
[0030] 1200 Processor
[0031] 1300 Memory
[0032] 1400 Other circuits. Detailed implementation manners
[0033] The present disclosure will be further described in detail below in conjunction with the drawings and embodiments. It can be understood that the specific implementation manners described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present disclosure are shown in the drawings.
[0034] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and embodiments.
[0035] Unless otherwise specified, the exemplary embodiments / examples shown will be understood to provide exemplary features of various details of some ways that can implement the technical concept of the present disclosure in practice. Therefore, unless otherwise specified, without departing from the technical concept of the present disclosure, the features of various embodiments / examples can be additionally combined, separated, interchanged, and / or rearranged.
[0036] The terms used in this disclosure are for the purpose of describing specific embodiments and are not restrictive. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. In addition, when the terms "comprising" and / or "including" and their variants are used in this specification, it is stated that there are the stated features, integers, steps, operations, components, assemblies, and / or groups thereof, but it does not exclude the presence or addition of one or more other features, integers, steps, operations, components, assemblies, and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about", and other similar terms are used as approximate terms rather than degree terms, so they are used to explain the inherent deviations of measured values, calculated values, and / or provided values that would be recognized by those of ordinary skill in the art.
[0037] Figure 1 It is a flowchart of a method for an anti-crawling station according to an exemplary embodiment of the present disclosure; Figure 2 It is a schematic diagram of the application of a method for an anti-crawling station according to an exemplary embodiment of the present disclosure; Figure 3 It is a schematic diagram of an array structure according to an exemplary embodiment of the present disclosure. The following will be combined with Figures 1 to 3 The specific implementation manners of each step of the above-mentioned method M100 for the anti-crawling station will be described in detail.
[0038] Step S102, determine a filtering rule for the address information in the red-black tree corresponding to the address hash value according to the address hash value used to characterize the address information in the access request.
[0039] Among them, a crawler is a program that crawls web page content for a search engine. The access request is generated by the crawler according to the search instruction given by the client to the search engine and is used to request the access permission of the website that meets the search instruction. The access request includes the network address information and domain name information of the crawler, etc. Usually, the crawling frequency of the crawler is relatively reasonable and the consumption of website resources is small; however, the web page crawling ability of malicious crawlers is poor, and they often send hundreds of requests concurrently and repeat crawling in a loop. These crawlers will cause a devastating blow to the website. Therefore, it is necessary to accurately identify malicious crawlers to maintain the order of the website and avoid unnecessary maintenance costs.
[0040] The address hash value is the result obtained by performing a hash operation on the address information in the access request and can characterize the address information in the access request. Among them, the hash operation is an algorithm that can accept address information of unlimited length and feedback a hash value of a fixed length to us. Different address information will obtain different hash values after the hash operation. Therefore, the address hash value has pertinence, uniqueness, and accuracy and can accurately characterize the address information of the crawler.
[0041] A red - black tree is a self - balancing binary search tree and a data structure for storing filtering rules. The root node of the red - black tree extends downward to two child nodes, and each child node further extends downward to two grand - child nodes, and so on. Among them, among the two child nodes with the same root node, the key value on the left is greater than the key value on the right. In the present disclosure, each node of the red - black tree stores a filtering rule, and the expiration time of each filtering rule is used as the key value of each node of the red - black tree. Among them, the filtering rules stored in the same red - black tree have the same rule characteristics.
[0042] The filtering rule stores response instructions corresponding to address information and multiple domain - name information. In other words, each filtering rule is a stored bitmap, which has multiple storage bits (for example, 30 * 8), and each storage bit stores the response instruction corresponding to one of the domain - name information of the same address information. Each filtering rule corresponds to an expiration time and is only valid when the expiration time is not reached.
[0043] In addition, the red - black tree is an important part of the constituent elements. Each element corresponds to a red - black tree and a doubly - linked list. Multiple elements form an array. In the array of the present disclosure, there are n + 1 elements, namely Bucket_0, Bucket_0 to Bucket_n. Of course, the number of elements can be set according to the actual situation and is not limited here. The construction method of the elements will be introduced later.
[0044] The specific implementation manner of step S102 is as follows: calculate the eigenvalue of the address hash value used to represent the address information in the access request to obtain the address feature corresponding to the address hash value; determine the first element corresponding to the address information according to the address feature, where the first element contains the filtering rules applicable to the address information; and based on the address hash value, traverse each filtering rule of the red - black tree in the first element to screen out the filtering rules adapted to the address information.
[0045] Among them, the eigenvalue calculation can be the remainder calculation of the address hash value, that is, divide the address hash value by the number of elements, and the part that is not divisible is the remainder, that is, the result of the remainder calculation, which is called the address feature in the present disclosure. For example, if the address hash value is 10 and the number of elements is 4, then the result of the eigenvalue calculation is 2, that is, the address feature is 2.
[0046] Each element in the array has a subscript representing the remainder result. If the number of elements is 4, then each element can be represented as Bucket_0, Bucket_1, Bucket_2, Bucket_3. At this time, the address information with an address feature of 2 corresponds to the element Bucket_2. In the present disclosure, the element corresponding to the address feature is defined as the first element, and then the red - black tree of the first element will contain the filtering rules applicable to this address information.
[0047] After determining the first element corresponding to the current address information, since each node of the red-black tree of the first element stores a filtering rule, and each filtering rule only has the same rule characteristics, which does not mean that the address hash values corresponding to each filtering rule are the same. Therefore, it is necessary to traverse each node in the red-black tree to filter out the filtering rules that match the address hash value to apply to the address information in the access request. Of course, the number of elements will be set according to the number of filtering rules. Therefore, there will not be too many filtering rules stored in the red-black tree of each element. For example, if there are one million filtering rules, 100,000 elements will be set, and each element contains 10 filtering rules. At this time, when traversing the filtering rules in any red-black tree, it will not cause an extension of the rule query time.
[0048] Multiple storage bits are set for different domain name information in the filtering rules. Therefore, after obtaining the filtering rule corresponding to the address information, it is also necessary to determine the storage bit of the current domain name information.
[0049] Step S104: Obtain the storage bit of the domain name information in the doubly linked list corresponding to the domain name hash value according to the domain name hash value used to represent the domain name information in the access request.
[0050] Among them, the domain name hash value is the result obtained by performing a hash operation on the domain name information in the access request and can represent the domain name information in the access request. Similarly, the hash operation is an algorithm that can accept domain name information of infinite length and feedback a hash value of a fixed length to us. Different domain name information will obtain different hash values after the hash operation. Therefore, the domain name hash value has pertinence, uniqueness, and accuracy and can accurately represent the domain name information of the crawler.
[0051] The doubly linked list is a chained storage structure that records the association relationship between each domain name information and the storage bit in the red-black tree. Usually, the storage bit of the domain name information in the red-black tree can be found according to the domain name information in the doubly linked list.
[0052] The specific implementation method of step S104 is: perform eigenvalue calculation on the domain name hash value used to represent the domain name information in the access request to obtain the domain name feature corresponding to the domain name information; determine the second element corresponding to the domain name feature according to the domain name feature, where the second element has a storage bit corresponding to the domain name feature; and based on the domain name hash value, traverse each storage bit of the doubly linked list in the second element to filter out the storage bit corresponding to the domain name feature.
[0053] Among them, the eigenvalue calculation can be the remainder calculation of the domain name hash value, that is, dividing the domain name hash value by the number of elements. The part that is not divisible is the remainder, that is, the result of the remainder calculation, which is called the domain name feature in this disclosure. For example, if the domain name hash value is 11 and the number of elements is 4, then the result of the eigenvalue calculation is 3, that is, the domain name feature is 3.
[0054] Referring to the above, the number of elements in the array is 4, and each element can be represented as Bucket_0, Bucket_1, Bucket_2, Bucket_3. At this time, the domain name information with a domain name feature of 3 corresponds to the element Bucket_3. In this disclosure, the element corresponding to the domain name feature is defined as the second element, and then the doubly linked list of the second element has storage bits corresponding to the domain name feature.
[0055] After determining the second element corresponding to the current domain name information, traverse each storage bit of the doubly linked list in the second element according to the domain name hash value, and filter out the storage bit corresponding to the domain name feature.
[0056] Step S106, filter out the response instruction corresponding to the storage bit in the filtering rule.
[0057] Among them, the response instruction is a strategy for adapting to the request body of the access request according to the address information and domain name information in the access request, and can be access prohibition, countermeasure operation, access permission, etc. For example, when the response instruction is 1, it means access prohibition or countermeasure operation; when the response instruction is 0, it means access permission, etc.
[0058] Specifically, map the storage bit corresponding to the domain name information to the filtering rule corresponding to the address information, and extract the response instruction corresponding to the current storage bit.
[0059] In some embodiments, before step S102, it further includes: performing a hash operation on the address information in the access request to obtain an address hash value for characterizing the address information in the access request.
[0060] In some embodiments, before step S102, it further includes: in response to the update time, respectively trigger multiple child processes to delete the expiration rules of the red-black trees in each element.
[0061] Among them, in order to facilitate the update and modification of the filtering rules in the red-black tree, an update time is set. When the update time is reached, the administrator can adjust each filtering rule in the red-black tree according to the current actual needs.
[0062] The timer is used for timing to accurately control the update time. The child process is a program that responds to the update time of the timer and deletes the expired rules in the red-black tree. To improve the deletion efficiency, the present disclosure provides multiple child processes, each child process is matched with multiple elements, and is only responsible for deleting the expired rules in the red-black tree of the elements that match it.
[0063] Specifically, according to the number of child processes and the number of elements, determine the multiple target elements matched by each child process; and in response to the update time, trigger each child process to delete the expired rules whose expiration time exceeds the target time according to the expiration time of each filtering rule in the red-black tree of the target elements.
[0064] Each child process has a corresponding subscript. For example, if the number of child processes is 8, then the subscripts of each child process can be sequentially set from 0 to 7. Divide the label of the element by the number of child processes, and the remainder is the part that cannot be divided, that is, the result of taking the remainder, which is called the element feature in the embodiment. For example, if the label of the element is 16 and the number of child processes is 8, then the element feature is 0, and at this time the element with the label of 16 corresponds to the child process with the subscript of 0. Traverse all elements by the above method, and then multiple corresponding elements can be matched for each child process. If there are 80 elements, by the above method, each child process can be matched with 10 elements.
[0065] Use the timer to determine the update time. For example, if the update time is 5 seconds, then the timer can be set to 5 seconds. When 5 seconds are reached, the timer triggers its corresponding child process. Although each child process corresponds to a timer respectively, the timing times of each timer are the same, and each child process executes the instruction to delete the expired rules simultaneously.
[0066] Based on the foregoing, the key values of each node in the red-black tree decrease sequentially from left to right, and the key value is assigned to the expiration time of each filtering rule. Therefore, the child process only needs to traverse the key values of the red-black tree sequentially from right to left until it identifies the key value whose expiration time is greater than the target time, and then delete the expired rules corresponding to all the key values on the right side of the key value, without needing to traverse the entire red-black tree. This can reduce unnecessary memory occupancy and also save the working time of the child process. Among them, the target time can be the current time or any set time.
[0067] In some embodiments, before step S102, it further includes: constructing multiple elements.
[0068] Specifically, a red-black tree and a doubly linked list are respectively constructed in each element; the rule hash values of each filtering rule are calculated; according to the rule hash values, the rule features of the filtering rules are obtained; according to the rule features, each filtering rule is assigned to the corresponding element; according to the expiration time of the filtering rules, the filtering rules in the same element are sequentially stored in the red-black tree.
[0069] When an administrator registers a filtering rule through the engine tool of the management server, the registration request may include the domain name of the management server, the address information of the client, multiple domain name information, penalty measures, and expiration time.
[0070] For example, the registration request is:
[0071] curl -H 'host:10.26.37.128-11000.zt.ke.com” http: / / 127.0.0.1:11000 / add?uuid=67e31090-0cfd-4209-aafd-ae0da0&punish=captcha_ban&expire=1617183514' -d 'pushapi.ke.com&10.26.37.128-90.zt.ke.com&www.abc.com&www.banana.com'
[0072] Among them,
[0073] curl -H 'host:10.26.37.128-11000.zt.ke.com” http: / / 127.0.0.1:11000 / is the domain name of the management server;
[0074] uuid=67e31090-0cfd-4209-aafd-ae0da0&punish=captcha_ban is the penalty measure for this filtering rule, countermeasure operation or prohibited access;
[0075] expire=1617183514'-d is the expiration time of this filtering rule;
[0076] pushapi.ke.com&10.26.37.128-90.zt.ke.com&www.abc.com&www.banana.com' are the multiple domain name information corresponding to this filtering rule.
[0077] Note that multiple domain name information in the registration request is separated by "&" so that after parsing each domain name information subsequently, a serial number can be defined for each domain name information. After storing each domain name information into the element, the storage bit corresponding to the serial number can be found in the storage bitmap according to the serial number, and the corresponding instruction of this storage bit is set to 1. Representing the registered domain name string with 1 storage bit can greatly reduce the memory usage, which is applicable to scenarios where a large number of filtering rules need to be stored.
[0078] After obtaining the registration request, the search engine server will parse the registration request to obtain the information related to the filtering rules in the registration request, including the address information of the client, domain name information, penalty measures, and expiration time.
[0079] Calculate the rule hash value of the filtering rule according to the above. The rule hash value is the result obtained by performing a hash operation on the filtering rule and can represent the filtering rule. Among them, the hash operation is an algorithm that can accept rule information of infinite length and feedback a hash value of a fixed length to us. Different filtering rules will obtain different hash values through the hash operation. Therefore, the rule hash value has pertinence, uniqueness, and accuracy, and can accurately represent different filtering rules.
[0080] The eigenvalue calculation can be the remainder calculation of the rule hash value, that is, dividing the rule hash value by the number of elements. The part that is not divisible is the remainder, that is, the result of the remainder calculation, which is called the rule feature in this embodiment. For example, if the rule hash value is 10 and the number of elements is 4, then the result of the eigenvalue calculation is 2, that is, the rule feature is 2. Based on the foregoing, if the elements are divided into Bucket_0, Bucket_1, Bucket_2, Bucket_3, then the filtering rule with a rule feature of 2 will be stored in the red-black tree of the element Bucket_2.
[0081] In this way, even if there are multiple filtering rules with the same rule hash value, multiple filtering rules can be stored into the same element at the same time.
[0082] In some embodiments, before step S104, it includes: performing a hash operation on the domain name information in the access request to obtain a domain name hash value for representing the domain name information in the access request.
[0083] In some embodiments, after step S106, it includes: in response to the coping instruction for executing the penalty measure, blocking the website scraping operation of the request subject of the access request. In response to the coping instruction for allowing access, allowing the website scraping operation of the request subject.
[0084] Query each element of the array using the engine tool according to the address information and domain name information in the crawler's access request. If the address information and domain name information are registered in the array, the crawler's access request is identified as a hacker request, and then a response instruction for implementing a penalty measure will be fed back to the crawler, and it is not allowed to crawl the web page content. The penalty measures can be either blocking access or countering the request subject, which can be selected according to requirements.
[0085] When both the address information and domain name information in the crawler's access request are not registered in the array, the crawler's access request is identified as a normal user request, and then a response instruction allowing access will be fed back to the crawler. At this time, it is allowed to crawl the web page content, and the business server proxies the access request.
[0086] According to the method for preventing website scraping of the present disclosure, the applicable filtering rule is determined in the red-black tree based on the address information, and then the storage position is determined in the doubly linked list based on the domain name information, and then the storage position is mapped to the filtering rule, so as to obtain the corresponding instruction corresponding to the address information and domain name information. Through the double verification of the address information and domain name information, the recognition accuracy of malicious crawlers is improved; by screening the response instructions through the red-black tree and the doubly linked list, the acquisition speed of the response instructions is improved.
[0087] Figure 4 It is a block diagram of a device for preventing website scraping according to an exemplary embodiment of the present disclosure. As Figure 4 shown, the present disclosure provides a device 1000 for preventing website scraping, including: a filtering rule determination module 1002, a storage position determination module 1004, and a response instruction generation module 1006. The filtering rule determination module 1002 is configured to determine the filtering rule for the address information in the red-black tree corresponding to the address hash value according to the address hash value used to represent the address information in the access request. The storage position determination module 1004 is configured to obtain the storage position of the domain name information in the doubly linked list corresponding to the domain name hash value according to the domain name hash value used to represent the domain name information in the access request. The response instruction generation module 1006 is configured to screen out the response instruction corresponding to the storage position in the filtering rule.
[0088] In some embodiments, the execution steps of the filtering rule determination module 1002 include: calculating the eigenvalue of the address hash value used to represent the address information in the access request to obtain the address feature corresponding to the address hash value; determining the first element corresponding to the address information according to the address feature, where the first element contains the filtering rule applicable to the address information; and traversing each filtering rule in the red-black tree of the first element based on the address hash value to screen out the filtering rule adapted to the address information.
[0089] In some embodiments, the execution steps of the storage bit determination module 1004 include: calculating eigenvalue of the domain name hash value for characterizing the domain name information in the access request to obtain the domain name feature corresponding to the domain name information; determining the second element corresponding to the domain name feature according to the domain name feature, wherein the second element has storage bits corresponding to the domain name feature; and traversing each storage bit of the doubly linked list in the second element based on the domain name hash value to filter out the storage bits corresponding to the domain name feature.
[0090] In some embodiments, it further includes an address hash value calculation module (not shown) for performing a hash operation on the address information in the access request to obtain an address hash value for characterizing the address information in the access request.
[0091] In some embodiments, it further includes an update module (not shown) for triggering multiple timers to delete the expiration rules of the red - black trees in each element in response to the update time.
[0092] In some embodiments, the execution steps of the update module (not shown) are: determining multiple target elements matched by each subprocess according to the number of subprocesses and the number of elements; and triggering each subprocess to delete the expiration rules whose expiration time exceeds the target time according to the expiration time of each filtering rule in the red - black tree of the target element in response to the update time.
[0093] In some embodiments, it further includes an element construction module (not shown) for constructing multiple elements, including: constructing a red - black tree and a doubly linked list in each element respectively; calculating the rule hash value of each filtering rule; obtaining the rule feature of the filtering rule according to the rule hash value; allocating each filtering rule to the corresponding element according to the rule feature; and storing the filtering rules in the same element into the red - black tree in sequence according to the expiration time of the filtering rule.
[0094] In some embodiments, it further includes a domain name hash value calculation module (not shown) for performing a hash operation on the domain name information in the access request to obtain a domain name hash value for characterizing the domain name information in the access request.
[0095] In some embodiments, it further includes an instruction response module (not shown) for blocking the crawling operation of the request subject of the access request in response to the response instruction for executing penalty measures; and allowing the crawling operation of the request subject in response to the response instruction for allowing access.
[0096] The device 1000 may include corresponding modules that execute each or several of the steps in the above flowchart. Therefore, each step or several steps in the above flowchart may be executed by corresponding modules, and the device may include one or more of these modules. The module may be one or more hardware modules specifically configured to execute the corresponding steps, or implemented by a processor configured to execute the corresponding steps, or stored in a computer-readable medium for implementation by the processor, or implemented through a certain combination.
[0097] The hardware structure may be implemented using a bus architecture. The bus architecture may include any number of interconnecting buses and bridges, depending on the specific application of the hardware and overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 may also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.
[0098] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Component (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.
[0099] For the device of the anti-crawling station according to the present disclosure, the applicable filtering rule is determined in the red-black tree through the address information, and then the storage location is determined in the doubly linked list through the domain name information, and then the storage location is mapped to the filtering rule, so as to obtain the corresponding instruction corresponding to the address information and the domain name information. Through the double verification of the address information and the domain name information, the recognition accuracy of malicious crawlers is improved; by screening the response instructions through the red-black tree and the doubly linked list, the acquisition speed of the response instructions is improved.
[0100] Any process or method description, whether represented in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code that includes one or more executable instructions for implementing a particular logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed in a substantially simultaneous manner or in an order opposite to that shown or discussed, depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain. A processor executes the various methods and processes described above. For example, the method embodiments in the present disclosure can be implemented as a software program tangibly embodied in a machine-readable medium, such as a memory. In some embodiments, part or all of the software program can be loaded and / or installed via the memory and / or a communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the processor can be configured to execute one of the above methods by any other suitable means (e.g., by means of firmware).
[0101] The logic and / or steps represented in a flowchart or otherwise described herein can be embodied in any readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch instructions from and execute instructions.
[0102] As used in this specification, a "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the readable storage medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the readable storage medium can even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a memory.
[0103] It should be understood that various parts of the present disclosure can be implemented by hardware, software, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0104] Those of ordinary skill in the art of this technology can understand that all or part of the steps of implementing the above embodiments of the method can be completed by a program instructing relevant hardware. The program can be stored in a readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0105] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. The storage medium can be a read-only memory, a disk, an optical disc, etc.
[0106] The present disclosure also provides an electronic device, including: a memory that stores execution instructions; and a processor or other hardware module that executes the execution instructions stored in the memory, so that the processor or other hardware module executes the above method.
[0107] The present disclosure also provides a readable storage medium in which execution instructions are stored. When the execution instructions are executed by a processor, they are used to implement the method of the anti-crawler station, including: determining a filtering rule for the address information in a red-black tree corresponding to the address hash value used to characterize the address information in the access request; obtaining the storage position of the domain name information in a doubly linked list corresponding to the domain name hash value used to characterize the domain name information in the access request; and screening out the response instruction corresponding to the storage position in the filtering rule.
[0108] The present disclosure also provides a computer program product, including computer programs / instructions that, when executed by a processor, implement the method of the anti-crawler station in any one of the embodiments of the present disclosure.
[0109] In the description of this specification, the descriptions referring to terms such as "one embodiment / way", "some embodiments / ways", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0110] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0111] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A method for preventing scratching stations, characterized in that, Including: Determine a filtering rule for the address information in the red - black tree corresponding to the address hash value used to characterize the address information in the access request; Obtain the storage location of the domain name information in the doubly - linked list corresponding to the domain name hash value used to characterize the domain name information in the access request; And Select, from the filtering rules, a response instruction corresponding to the storage location. The response instruction is a strategy for adaptively dealing with the request body of the access request according to the address information and domain name information in the access request.
2. The method of the anti-grabbing station according to claim 1, characterized in that, The step of determining a filtering rule for the address information in the red - black tree corresponding to the address hash value used to characterize the address information in the access request includes: Perform eigenvalue calculation on the address hash value used to characterize the address information in the access request to obtain an address feature corresponding to the address hash value; Determine a first element corresponding to the address information according to the address feature, where the first element contains a filtering rule applicable to the address information; and Based on the address hash value, traverse each filtering rule in the red - black tree of the first element and select a filtering rule adapted to the address information.
3. The method of the anti-grabbing station according to claim 1, characterized in that, The step of obtaining the storage location of the domain name information in the doubly - linked list corresponding to the domain name hash value used to characterize the domain name information in the access request includes: Perform eigenvalue calculation on the domain name hash value used to characterize the domain name information in the access request to obtain a domain name feature corresponding to the domain name information; Determine a second element corresponding to the domain name feature according to the domain name feature, where the second element has a storage location corresponding to the domain name feature; and Based on the domain name hash value, traverse each storage location in the doubly - linked list of the second element and select a storage location corresponding to the domain name feature.
4. The method of the anti-grabbing station according to claim 1, characterized in that, Before the step of determining a filtering rule for the address information in the red - black tree corresponding to the address hash value used to characterize the address information in the access request, it further includes: Perform a hash operation on the address information in the access request to obtain an address hash value used to characterize the address information in the access request.
5. The method of the anti-grabbing station according to claim 1, characterized in that Before the step of determining a filtering rule for the address information in the red - black tree corresponding to the address hash value used to characterize the address information in the access request, it further includes: In response to the update time, trigger multiple timers respectively to delete the expiration rules in the red - black trees of each element.
6. The method of the anti-grabbing station according to claim 5, characterized in that The step of, in response to the update time, triggering multiple subprocesses to delete the expiration rules in the red - black trees of each element includes: Determine multiple target elements matched by each subprocess according to the number of subprocesses and the number of elements; and In response to the update time, trigger each subprocess respectively to delete the expiration rules whose expiration time exceeds the target time according to the expiration time of each filtering rule in the red - black tree of the target element.
7. The method of the anti-grabbing station according to claim 1, characterized in that Before determining a filtering rule for the address information in a red - black tree corresponding to the address hash value used to characterize the address information in the access request, it includes: Construct multiple elements, including: Construct a red - black tree and a doubly - linked list in each of the elements respectively; Calculate the rule hash value of each of the filtering rules; Obtain the rule feature of the filtering rule according to the rule hash value; Allocate each of the filtering rules to the corresponding element according to the rule feature; According to the expiration time of the filtering rule, store the filtering rules in the same element sequentially into the red - black tree.
8. The method of the anti-grabbing station according to claim 1, wherein Before obtaining the storage location of the domain name information in a doubly - linked list corresponding to the domain name hash value used to characterize the domain name information in the access request, it includes: Perform a hash operation on the domain name information in the access request to obtain a domain name hash value used to characterize the domain name information in the access request.
9. The method of the anti-grabbing station according to claim 1, wherein After screening out the response instruction corresponding to the storage location from the filtering rules, it includes: In response to the response instruction for executing a penalty measure, prevent the website scraping operation of the request subject of the access request; and In response to the response instruction for allowing access, allow the website scraping operation of the request subject.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the anti - website - scraping method according to any one of claims 1 to 9.
11. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and the computer program is suitable for being loaded by the processor to execute the anti - website - scraping method according to any one of claims 1 to 9.
12. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the anti - website - scraping method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Website access control method and device based on DNS analysis
CN111314301A
Data processing method, apparatus and device, and computer readable storage medium
CN114466054A