A crawler-based IP proxy pool management method and system
By identifying the risk control type of the target website and dynamically adjusting the proxy IP strategy, the problem of inappropriate IP type selection and high cost in existing technologies is solved. This enables efficient and low-cost data crawling, supports real-time monitoring and adjustment, and adapts to changes in the risk control strategy of the target website.
Patent Information
- Application Number
- CN202510119260.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing IP proxy pool management technologies cannot flexibly cope with the risk control strategies and crawler task requirements of different target websites, resulting in inappropriate IP type selection, high costs, waste of resources and high labor costs, and a lack of automatic dynamic adjustment capabilities.
By identifying the risk control type of the target website, dynamically adjusting the proxy IP strategy, generating IP strategies based on the characteristics of the crawling task, optimizing IP procurement, and monitoring the crawling process in real time for IP allocation and recycling, intelligent and low-cost data crawling is achieved.
It implements an anti-crawling mechanism that accurately identifies target websites, dynamically adjusts proxy IP strategies, improves crawling success rate, avoids resource waste, reduces costs, quickly responds to changes in risk control strategies, and ensures continuous low-cost operation of crawling tasks.
Smart Images

Figure CN120179888B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of web crawlers and IP proxy, in particular to a crawler-based IP proxy pool management method and system. BACKGROUND
[0002] In network data collection, different target websites have different risk control systems and anti-crawling mechanisms. For example, some websites will trigger a ban mechanism if the same account frequently changes IP, while other websites need to switch IP frequently to avoid detection; traditional IP proxy pool technology usually lacks precise judgment of the risk control mechanism of the target website. In network data collection, the selection of proxy IP is an important factor to ensure the smooth completion of crawling tasks. Existing technologies mainly focus on the use of the following four types of proxy IP:
[0003] The first type is tunnel IP, which automatically changes IP each time a request is made. This type of IP is suitable for short-term concurrent crawling, but for tasks that require login or long-term crawling, its frequent switching will be recognized as abnormal behavior by the target website's risk control system, resulting in account bans or increased verification frequency, thereby increasing the cost of collection.
[0004] The second type is time-sensitive IP, which has a fixed validity period (such as 30 minutes, 1 hour, 12 hours, etc.). This type of IP can be reused in the short term, but when the task duration exceeds the IP validity period, it needs to be frequently switched to a new IP, resulting in additional costs. However, if the same website and account are continuously collected for 30 minutes, 1 hour, or 12 hours, it will be banned. This type of IP is the lowest cost option, but if you choose a 1-hour IP for collection, the account can be used for 12 hours, which will increase the cost of the account.
[0005] The third type is fixed IP, which has a long validity period (such as 24 hours, 1 week, or 1 month). According to the risk control standards of the target website, this type is suitable for tasks that do not require IP detection but require long-term login collection or large amounts of collection. Selecting this type of IP is the most cost-effective option.
[0006] The fourth type is other customized types of IP (overseas IP, real residential IP), etc.
[0007] Due to significant differences in risk control systems and crawler task requirements for different target websites, how to select the optimal IP type according to the actual scenario becomes critical. Existing proxy pool technologies usually cannot flexibly respond to these demand differences, resulting in the following problems:
[0008] (1) Inappropriate IP type selection: using tunnel IP for tasks that require long-term stable IP, resulting in account bans or frequent verification; conversely, using high-cost fixed IP for short-term tasks results in resource waste.
[0009] (2) High cost: unable to automatically optimize IP procurement strategy according to the actual needs of the target website, resulting in a large number of procurement and task mismatched proxy IP resources.
[0010] (3) Lack of flexibility: lack of automatic dynamic adjustment capability, unable to respond to changes in the target website's risk control strategy.
[0011] (4) High labor cost and maintenance cost: traditional proxy IP pool requires human intervention, manual observation and recording of website risk control frequency, and finally analysis and debugging to determine the type of IP to be purchased. In the face of multiple and miscellaneous collection channels, a large number of manual intervention is required.
[0012] Therefore, the existing IP proxy pool management technology cannot meet the needs of different target website risk control strategies and crawler tasks, resulting in deficiencies in intelligence, efficiency, cost savings, and flexibility. SUMMARY
[0013] To solve the above problems, the present disclosure proposes a crawler-based IP proxy pool management method and system, which identifies the risk control type of the target website, dynamically adjusts the proxy IP strategy, and efficiently executes the data crawling task on the target website.
[0014] According to some embodiments, the present disclosure adopts the following technical solutions:
[0015] A crawler-based IP proxy pool management method, comprising:
[0016] Obtaining feedback data of the target website on the crawler crawling, analyzing the anti-crawling mechanism of the target website based on the feedback data, and identifying the risk control type of the target website;
[0017] Combining the characteristics of the to-be-executed crawling task and the risk control type of the target website, generating an IP strategy for the task;
[0018] According to the IP strategy of the task, optimizing the cost calculation, purchasing IP for the to-be-executed crawling task, and adding it to the IP proxy pool;
[0019] The IP is used as a proxy IP for the to-be-executed crawling task to execute data crawling on the target website, and the crawling process is monitored in real time. According to the changes in the anti-crawling mechanism of the target website, the IP for the crawling task is switched, and the dynamic allocation and recycling of the IP in the IP proxy pool are realized.
[0020] According to some embodiments, the present disclosure adopts the following technical solutions:
[0021] A crawler-based IP proxy pool management system, comprising:
[0022] The risk control identification module is configured to: acquire feedback data of the target website to the crawler, analyze the anti-crawler mechanism of the target website based on the feedback data, and further identify the risk control type of the target website.
[0023] The policy generation module is configured to: generate an IP policy of the task in combination with the characteristics of the to-be-executed crawling task and the risk control type of the target website.
[0024] The IP procurement module is configured to: procure IP for the to-be-executed crawling task according to the IP policy of the task through cost calculation optimization, and add the IP to an IP proxy pool.
[0025] The task execution module is configured to: execute data crawling of the target website by taking the IP as a proxy IP of the to-be-executed crawling task, monitor the crawling process in real time, switch the IP for the crawling task according to the change of the anti-crawler mechanism of the target website, and realize dynamic allocation and recycling of the IP in the IP proxy pool.
[0026] According to some embodiments, the present disclosure adopts the technical scheme as follows:
[0027] A computer program product comprising a computer program, which, when executed by a processor, implements the crawler-based IP proxy pool management method.
[0028] According to some embodiments, the present disclosure adopts the technical scheme as follows:
[0029] A non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implements the crawler-based IP proxy pool management method.
[0030] According to some embodiments, the present disclosure adopts the technical scheme as follows:
[0031] An electronic device comprising a processor, a memory, and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the crawler-based IP proxy pool management method.
[0032] Compared with the prior art, the present disclosure has the following beneficial effects:
[0033] The present disclosure provides a crawler-based IP proxy pool management method and system, which identifies the risk control type of a target website, dynamically adjusts a proxy IP policy, and efficiently executes a data crawling task of the target website. The specific advantages are as follows:
[0034] 1. Intelligence: Accurately identify the anti-crawling mechanism and risk control type of the target website, dynamically adjust the proxy IP strategy, and achieve high success rate of the task.
[0035] 2. Efficiency: Through the precise matching of task characteristics and IP types, waste of IP resources is avoided, and crawling efficiency is improved.
[0036] 3. Cost savings: Preferentially use free or low-cost IP, reduce the frequency of using paid IP, and avoid additional costs caused by excessive IP switching.
[0037] 4. Flexibility: Supports real-time monitoring and adjustment, can quickly respond to changes in target website risk control strategies, and ensures the continuous and low-cost operation of the crawler task. BRIEF DESCRIPTION OF DRAWINGS
[0038] The accompanying drawings, which are part of this disclosure, serve to provide further understanding of the present disclosure, and the illustrative embodiments of the present disclosure and their descriptions serve to explain the present disclosure, and do not constitute an improper limitation on the present disclosure.
[0039] Figure 1 A flowchart of a crawler-based IP proxy pool management method according to Embodiment 1. DETAILED DESCRIPTION
[0040] The present disclosure will be further described below in conjunction with the drawings and embodiments.
[0041] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.
[0042] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form, and in addition, it should be understood that when the terms "comprise" and / or "comprise" are used in the specification, they indicate the presence of a feature, step, operation, device, component, and / or combination thereof.
[0043] Embodiment 1
[0044] In an embodiment of the present disclosure, a crawler-based IP proxy pool management method is provided, comprising:
[0045] Step S1: Obtain feedback data of the target website to the crawler crawling, analyze the anti-crawling mechanism of the target website based on the feedback data, and further identify the risk control type of the target website;
[0046] Step S2: Generate IP strategy for the task based on the characteristics of the crawling task to be performed and the risk control type of the target website;
[0047] Step S3: According to the IP strategy of the task, optimize the cost calculation, purchase IP for the crawling task to be performed, and add it to the IP proxy pool;
[0048] Step S4: Use the IP as a proxy IP for the crawling task to be performed, perform data crawling on the target website, and monitor the crawling process in real time. According to the changes in the anti-crawling mechanism of the target website, switch the IP for the crawling task, and realize dynamic allocation and recycling of IP in the IP proxy pool.
[0049] As an embodiment, the IP proxy pool management method based on crawler of the present disclosure dynamically adjusts, orders or selects the optimal proxy IP according to the anti-crawling mechanism and risk control type of the target website, so as to realize efficient and low-cost network data collection, such as Figure 1 As shown in the figure, the specific implementation process is as follows:
[0050] Step 1: Target website risk identification
[0051] Step 1.1: Automatically analyze the anti-crawling mechanism and risk control type of the target website, analyze IP switching sensitivity, access frequency limit, account behavior monitoring rules and other characteristics.
[0052] Specifically, the crawler system accesses the login page and data interface of the target website to obtain feedback data of the website (including ban prompt, verification code trigger, access limit), and analyzes the following risk control characteristics, i.e. Anti-crawling mechanism, specifically:
[0053] IP switching sensitivity: Whether IP switching triggers a verification code or a ban;
[0054] Access frequency limit: Whether the same account needs stable IP support;
[0055] Account behavior monitoring rules: Whether the access frequency limit is associated with the account or IP.
[0056] Step 1.2: Based on the analysis result, classify the target website into frequent switching type, stable IP type and mixed type, etc. so as to generate targeted IP strategy in step 2.
[0057] Specifically, the classification characteristics corresponding to the three types of risk control are obtained by collecting feedback data of the target website, including access log, response time, verification code trigger frequency, ban frequency, website response time change, bandwidth change, proxy IP request success rate, request failure state reason and other log files, through analysis and classification, specifically:
[0058] (1) Frequent switching type:
[0059] IP switching sensitivity: If the website does not trigger a verification code or ban for multiple accesses to the same account from different IPs within a short period of time (300 seconds, 300 requests), it indicates that the website is not sensitive to IP switching.
[0060] Access frequency sensitivity: If the website allows high-frequency (switching IP every request, 60 seconds > 180 times) IP switching without limiting the account usage state, it means that the account and IP binding are not subject to risk control calculations, and the frequently switched IP can be selected to evade website verification codes or other risk control issues.
[0061] Account behavior monitoring rules: If the website does not strictly monitor the access behavior of the account, it allows the same account to frequently log in or request data within a short period of time (60 seconds).
[0062] Identification method: By simulating frequent IP switching requests, observe whether the risk control measures are triggered. If there is no obvious restriction or ban, it can be judged as frequently switched.
[0063] (2) Stable IP type:
[0064] IP switching sensitivity: If the website immediately bans or frequently triggers a verification code (60 requests per minute, once a verification code appears) after IP switching or multiple (more than 10 times within 2 hours) IP switching, it indicates that it is very sensitive to IP switching.
[0065] Access frequency limit: If the website requires a high request frequency (less than 1 per second) for access frequency, and the same IP needs to be connected for a long time (more than 6 hours).
[0066] Account behavior monitoring rules: If the website has strict monitoring of the access behavior of the account, it requires a stable IP support or says that the longer the same IP remains logged in, the lower the frequency of triggering risk control.
[0067] Identification method: By maintaining the same IP for a long time (more than 12 hours) access, analyze whether the log reduces the triggering of the verification code or ban. If stable access does not trigger risk control measures or significantly reduces the frequency / cost of risk control compared to other types of IP, it can be judged as a stable IP type.
[0068] (3) Mixed type:
[0069] IP switching sensitivity: The website has a certain tolerance for IP and account switching, but may trigger risk control measures when frequently switching.
[0070] Access frequency limit: The website allows a certain frequency of access, but may limit requests when reaching a certain threshold.
[0071] Account behavior monitoring rules: The website has certain monitoring of the access behavior of the account, but allows flexible operation under certain conditions, such as the website's search interface does not allow you to make high-concurrency requests or switch IPs, but some detail lists allow high-concurrency frequent account switching for collection
[0072] Identification method: By testing different IP switching frequencies and access frequencies, observe the triggering of the risk control measures. If flexible operation is allowed under certain conditions, but triggers risk control under other conditions, it can be judged as mixed type. Such as the number of successful times, the number of times triggering risk control, the cost of solving risk control, the cost of frequent IP and stable IP are compared to judge comprehensively.
[0073] For example, according to the analysis result, the target website is classified as "stable IP type", and step 2 generates the corresponding IP strategy: prefer to use fixed IP or long-acting time-sensitive IP.
[0074] Step 2: Task IP strategy generation
[0075] According to the specific crawling task demand characteristics (whether to log in, collection duration, data size) and the risk control type of the target website, generate the IP strategy of the specific task.
[0076] Provide collection strategy optimization for different tasks such as single-account multi-thread login, multi-account high-concurrency collection, and long-time stable collection. The specific collection method can be analyzed by combining the crawler collection end, such as judging that A website can be collected using IP frequency type, which makes multiple requests or changes IP in a short time. If the target website does not control the account, the crawler collection end can definitely use multi-account high-concurrency collection to save time cost, and single-account multi-threading may be suitable for IP stable type, which is long-time request with low frequency. Long-time stable collection is generally a mixed IP scenario, such as the website has different risk control mechanisms for IP on different interfaces, which requires long-time low request to complete the task.
[0077] For example, assume that a crawling task needs to log in to an e-commerce platform for a long time and collect product information, with the following demand characteristics:
[0078] (1) Single-account collection duration is about 12 hours;
[0079] (2) Large data volume, need stable network environment.
[0080] Then, combined with the demand characteristics and the risk control type of the target website, the following IP strategy is generated:
[0081] (1) Use 12-hour or 24-hour fixed IP;
[0082] (2) Configure a backup IP, switch 5 minutes in advance to avoid task interruption due to IP failure.
[0083] Step 3: Intelligent IP Procurement
[0084] According to the IP strategy, automatically select IP channel merchants, and automatically order the most suitable IP types from agent service providers, including tunnel IP, time-limited IP (30 minutes to 12 hours), fixed IP (24 hours or more), and other customized IP types. Optimize procurement logic to prioritize the use of free or low-cost IP resources and call high-quality paid IP as needed.
[0085] For example, according to the above IP strategy, order the following IP from the agent service provider:
[0086] 1 24-hour fixed IP for main tasks;
[0087] 1 standby 12-hour time-limited IP.
[0088] The procurement process is optimized through cost calculation: prioritize lower-priced IP resources to avoid wasting high-quality IP resources.
[0089] Step 4: IP Allocation and Switching
[0090] Distribute the purchased proxy IP to specific crawler tasks according to the strategy. For long-term collection tasks, predict the IP usage expiration time in advance and dynamically switch to standby IP to ensure uninterrupted tasks.
[0091] During the crawling process, monitor the target website feedback data in real time, and then monitor the changes in anti-crawling mechanisms:
[0092] Response data includes blocking prompts, CAPTCHA triggers, access restrictions, and other behaviors. By analyzing the monitored response data, analyze the changes in anti-crawling mechanisms, and adjust the current task's IP strategy based on the monitoring data analysis results, such as switching IP types or changing the switching frequency to adapt to changes in the target website's risk control strategy.
[0093] For example, the target website suddenly increases access restrictions on accounts, triggering a ban after exceeding a certain threshold number of logins per hour. After real-time monitoring of this change, the IP strategy is adjusted as follows:
[0094] (1) Switch to standby IP to reduce access frequency;
[0095] (2) For subsequent tasks targeting this website, adjust the strategy to prioritize the use of longer-term fixed IP to avoid frequent switching.
[0096] Step 5: Multi-task Concurrency Crawling
[0097] For tasks that crawl multiple target websites simultaneously, generate independent IP strategies based on the risk control type and task requirements of each website, for example:
[0098] Website A is a frequent switching type, using tunnel IP for high concurrency crawling;
[0099] Website B is a stable IP type, using 24-hour fixed IP to complete login collection.
[0100] According to the progress of the task and the resource usage, the IP is dynamically allocated and recycled, and the utilization rate of the IP pool is improved.
[0101] Embodiment 2
[0102] In an embodiment of the present disclosure, a crawler-based IP proxy pool management system is provided, comprising:
[0103] The risk control recognition module is configured to: obtain feedback data of the target website to the crawler crawling, analyze the anti-crawling mechanism of the target website based on the feedback data, and further recognize the risk control type of the target website;
[0104] The strategy generation module is configured to: generate the IP strategy of the task in combination with the characteristics of the to-be-executed crawling task and the risk control type of the target website;
[0105] The IP procurement module is configured to: according to the IP strategy of the task, through cost calculation optimization, procure IP for the to-be-executed crawling task, and add the IP to the IP proxy pool;
[0106] The task execution module is configured to: use the IP as the proxy IP of the to-be-executed crawling task, execute data crawling on the target website, monitor the crawling process in real time, switch the IP for the crawling task according to the change of the anti-crawling mechanism of the target website, and realize dynamic allocation and recycling of the IP in the IP proxy pool.
[0107] Embodiment 3
[0108] In an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the crawler-based IP proxy pool management method.
[0109] Embodiment 4
[0110] In an embodiment of the present disclosure, a non-transitory computer readable storage medium is provided, which is used to store computer instructions, and the computer instructions, when executed by a processor, implement the crawler-based IP proxy pool management method.
[0111] Embodiment 5
[0112] An embodiment of the present disclosure provides an electronic device, comprising: a processor, a memory and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for managing an IP proxy pool based on a crawler.
[0113] The present disclosure is described with reference to the flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The means for implementing the functions specified in one block or multiple blocks.
[0114] These computer program instructions can also be loaded into a computer or other programmable data processing device to cause a series of operation steps to be performed on the computer or other programmable data processing device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The steps for implementing the functions specified in one block or multiple blocks.
[0115] Although the specific embodiments of the present disclosure are described above with reference to the accompanying drawings, the present disclosure is not limited to the above embodiments, and various modifications or changes can be made to the embodiments without departing from the scope of the present disclosure.
Claims
1. A crawler-based IP proxy pool management method, characterized in that, The method comprises: Obtaining feedback data of the target website to the crawler, analyzing the anti-crawling mechanism of the target website based on the feedback data, and identifying the risk control type of the target website; the risk control type includes frequent switching type, stable IP type and mixed type; Generating an IP strategy for the task in combination with the characteristics of the to-be-executed crawling task and the risk control type of the target website; According to the IP strategy of the task, the IP for the to-be-executed crawling task is purchased through cost calculation optimization and added to the IP proxy pool; the IP for the to-be-executed crawling task is purchased, specifically: According to the IP strategy, an IP channel merchant is automatically selected, and an IP type corresponding to the proxy service provider is ordered; Optimize the purchase logic, preferentially use free or low-cost IP resources, and call high-quality paid IP as needed; The IP is used as a proxy IP for the to-be-executed crawling task to execute data crawling on the target website, and the crawling process is monitored in real time, the IP is switched for the crawling task according to the change of the anti-crawling mechanism of the target website, and dynamic allocation and recycling of the IP in the IP proxy pool are realized.
2. The crawler-based IP proxy pool management method of claim 1, wherein, The anti-crawling mechanism of the target website based on the feedback data is analyzed, specifically: IP switching sensitivity: whether IP switching triggers a verification code or a ban; Access frequency limit: whether the same account needs stable IP support; Account behavior monitoring rules: whether the access frequency limit is associated with the account or the IP.
3. The method of claim 1, wherein the method further comprises: The characteristics of the to-be-executed crawling task include whether to log in, collection duration, and data size.
4. The crawler-based IP proxy pool management method of claim 1, wherein, The IP strategy includes single-account multi-threaded login, multi-account high-concurrency collection, and long-time stable collection.
5. A crawler-based IP proxy pool management system, characterized in that, The method comprises: The risk control identification module is configured to: obtain feedback data of the target website to the crawler, analyze the anti-crawling mechanism of the target website based on the feedback data, and identify the risk control type of the target website; the risk control type includes frequent switching type, stable IP type and mixed type; The strategy generation module is configured to: generate an IP strategy for the task in combination with the characteristics of the to-be-executed crawling task and the risk control type of the target website; The IP procurement module is configured to: according to the IP strategy of the task, the IP for the to-be-executed crawling task is purchased through cost calculation optimization and added to the IP proxy pool; the IP for the to-be-executed crawling task is purchased, specifically: According to the IP strategy, an IP channel merchant is automatically selected, and an IP type corresponding to the proxy service provider is ordered; Optimize the purchase logic, preferentially use free or low-cost IP resources, and call high-quality paid IP as needed; The task execution module is configured to: use the IP as a proxy IP for the to-be-executed crawling task to execute data crawling on the target website, and monitor the crawling process in real time, switch the IP for the crawling task according to the change of the anti-crawling mechanism of the target website, and realize dynamic allocation and recycling of the IP in the IP proxy pool.
6. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the IP proxy pool management method based on the crawler in any one of claims 1-4.
7. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium is used to store computer instructions, and the computer instructions are executed by the processor to realize the IP proxy pool management method based on the crawler in any one of claims 1-4.
8. An electronic device, comprising: The method comprises: A processor, a memory and a computer program; wherein the processor is connected with the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for managing the IP proxy pool based on the crawler as claimed in any one of claims 1-4.
Citation Information
Patent Citations
Multi-website parallel crawling IP agent pool construction system and method
CN114143290A
Cloud financial data acquisition method based on web crawler technology
CN117474694A