Multi-channel-based web page data crawling method, system, device and medium
Patent Information
- Application Number
- CN202410713553.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-04
AI Technical Summary
[0004]本发明旨在至少在一定程度上解决现有技术中的技术问题之一,通过对数据爬取进行改进,用于解决现有技术中因缺少对使用多个代理节点对网页数据进行爬取的分析,在单一节点爬取网页数据的情况下,为了避免过度频繁或长时间快速访问网页从而被限制,爬虫程序的爬取频率设置过低,导致爬取效率低下的问题
[0048]本发明的有益效果:本发明通过对代理节点进行划分,得到测试节点,再通过频率测试循环按爬取频率从小到大的顺序进行数据爬取,再分析数据爬取结果,从而得到目标网页的频率阈值,这样的好处在于,通过较少的测试节点能够快速地测试得到频率阈值,同时测试时的爬取频率按递增的方式进行变化,能够以最少的IP地址被限制为代价,得到较为准确的频率阈值,避免了IP资源的浪费,同时提高了阈值检测的高效性;
Smart Images

Figure CN118467807B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data crawling technology, specifically to a method, system, device, and medium for web page data crawling based on multiple channels. Background Technology
[0002] Data crawling, also known as web crawling or web scraping, refers to the process of simulating a browser to send network requests and receive response requests. It is an automated process that obtains and extracts data from the Internet according to certain rules. It can be used to collect, parse, and store information from sources such as web pages, API interfaces, and documents.
[0003] Existing improvements to data crawling typically focus on enhancing the crawling method within a single-channel context, aiming to make it more intelligent and efficient. For example, Chinese patent CN116186368A discloses a data crawling method and system. This approach involves creating an initial keyword list, sorting it according to weight ratios, and allocating keywords within the initial keyword list based on a sorting algorithm. The crawler then crawls data based on these keywords. However, existing improvements lack analysis of crawling webpage data through multiple proxy nodes. When crawling webpage data from a single node, the crawling frequency is often set too low to avoid excessively frequent or prolonged rapid access to the webpage, resulting in low crawling efficiency. Therefore, it is necessary to improve existing data crawling methods. Summary of the Invention
[0004] This invention aims to at least partially solve one of the technical problems in the prior art. By improving data crawling, it addresses the issue that the prior art lacks analysis on crawling web page data using multiple proxy nodes. In the case of crawling web page data with a single node, the crawling frequency of the crawler program is set too low to avoid being restricted due to excessively frequent or long-term rapid access to web pages, resulting in low crawling efficiency.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a multi-channel web page data crawling method, comprising:
[0006] Get the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output the test nodes and crawl nodes;
[0007] Set different crawling frequencies for test nodes, crawl data from the target website, and output multiple crawling results; analyze the multiple crawling results and output frequency thresholds based on the analysis results;
[0008] The crawling nodes are grouped and output as multiple node groups, and the node groups are numbered; the crawling frequency of each crawling node is set based on a frequency threshold, and the crawling nodes in the node group are controlled to crawl concurrently based on the node group number; crawling information or reconstructed node information is output, and regrouping is performed based on the reconstructed node information.
[0009] The crawled information is integrated and the page information is output.
[0010] Furthermore, obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output the test nodes and crawling nodes, including:
[0011] Get the total number of proxy nodes and mark it as the total number of nodes;
[0012] Get the maximum crawling frequency of each proxy node and mark it as a node at full frequency;
[0013] Sort the proxy nodes according to their full frequency to obtain the node sequence;
[0014] Each proxy node is numbered in ascending order from left to right. Proxy nodes with numbers greater than the total number of nodes with the first scaling factor are marked as test nodes, and proxy nodes with numbers less than or equal to the total number of nodes with the first scaling factor are marked as crawling nodes.
[0015] Furthermore, different crawling frequencies are set for the test nodes to crawl data from the target website, outputting multiple crawling results; these results are analyzed, and frequency thresholds are output based on the analysis results, including:
[0016] Set the smallest number in the test node to i, and set n to an initial value of 0;
[0017] Execute a frequency test loop, the frequency test loop including:
[0018] The crawling frequency of the test node numbered i+n is set to h×(n+1), where h is the first frequency. The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, no processing is performed; when the crawling fails, a failure message is output.
[0019] If no failure message is received, increment n by one and execute the frequency test loop again.
[0020] When a failure message is received, n is decremented by one, and the crawling frequency of the test node numbered i+n is set to h×(n+1.5). The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, the crawling frequency is marked as the frequency threshold; when the crawling fails, h×(n+1) is set as the frequency threshold.
[0021] End the frequency test loop;
[0022] Mark the test nodes with numbers less than i+n back as crawling nodes.
[0023] Furthermore, the crawled nodes are grouped, resulting in multiple node groups. These node groups are then numbered, including:
[0024] Execute a group processing loop, the group processing loop including:
[0025] Set m to an initial value of 1;
[0026] The first random number of crawling nodes are assigned to the same group, which is marked as a node group and numbered m.
[0027] The second random number of crawling nodes are assigned to the same group, and the node group number is m×2;
[0028] The third random number of crawling nodes are assigned to the same group, and the node group is numbered m×3;
[0029] Determine if there are any remaining crawlable nodes. If the result is yes, perform a quantity check; if the result is no, end the group processing loop.
[0030] The quantity judgment process includes: judging whether the number of remaining crawling nodes is greater than the third number of the second multiple; if the judgment result is yes, increment m by one and execute the grouping processing loop again; if the judgment result is no, divide the remaining crawling nodes into the same group and number the node group as (m×3)+1.
[0031] Furthermore, the crawling frequency of each crawling node is set based on a frequency threshold, and the concurrent crawling of crawling nodes in a node group is controlled based on the node group number, including:
[0032] Set the crawling frequency of all crawled nodes to the frequency threshold of the second scaling factor;
[0033] Execute a concurrent crawling loop, the concurrent crawling loop including:
[0034] Reset m to 1 and n to 0. Control all proxy nodes in the node group with node number m to simultaneously crawl web page data using the crawler program according to the crawling frequency. When the crawl is successful, output the crawling information and assign the crawling information number n. Increment n and execute the completion judgment. When the crawl fails, end the concurrent crawling loop and output the recombined node information.
[0035] The completion judgment process includes: judging whether the web page data has been completely crawled; when the judgment result is yes, outputting crawling completion information and ending the concurrent crawling loop; when the judgment result is no, incrementing m by one and executing the concurrent crawling loop again.
[0036] Furthermore, monitoring the crawling results and performing regrouping based on the monitoring results includes:
[0037] When reorganization node information is received, delete all node groups and scale the first, second, and third quantities by the second scaling factor.
[0038] Repeat the group processing loop and the concurrent crawling loop.
[0039] Furthermore, the crawled information is integrated, and the output page information includes:
[0040] Sort all crawled information in ascending order of their numbers, mark all sorted crawled information as page information, and output the page information.
[0041] Secondly, the present invention also provides a web page data crawling system based on multiple channels, including a threshold testing module, a node grouping module, and a concurrent crawling module; the threshold testing module is used to obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output test nodes and crawling nodes.
[0042] It is also used to set different crawling frequencies for test nodes, crawl data from target websites, and output multiple crawling results; analyze multiple crawling results, and output frequency thresholds based on the analysis results;
[0043] The node grouping module is used to group crawled nodes, output multiple node groups, and number the node groups.
[0044] The concurrent crawling module is used to set the crawling frequency of each crawling node based on a frequency threshold, control the concurrent crawling of crawling nodes in the node group based on the node group number, output crawling information or reassemble node information, and perform regrouping processing based on the reassembled node information.
[0045] It is also used to integrate crawled information and output page information.
[0046] Thirdly, this application provides an electronic device including a processor and a memory, the memory storing computer-readable instructions, which, when executed by the processor, perform the steps of the method described above.
[0047] Fourthly, this application provides a storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the method described above.
[0048] The beneficial effects of this invention are as follows: This invention divides proxy nodes to obtain test nodes, and then crawls data in ascending order of crawling frequency through frequency testing loops. The data crawling results are then analyzed to obtain the frequency threshold of the target webpage. The advantage of this is that the frequency threshold can be obtained quickly with fewer test nodes. At the same time, the crawling frequency during the test changes in an increasing manner, which can obtain a more accurate frequency threshold at the cost of limiting the fewest IP addresses, avoiding the waste of IP resources and improving the efficiency of threshold detection.
[0049] This invention also groups the crawling nodes, numbers the groups, sets the crawling frequency of each crawling node based on a frequency threshold, and finally controls the crawling nodes to crawl the target webpage according to the crawling frequency to obtain the final page data. The advantage of this is that there are three different numbers of node groups, which are crawled concurrently in sequence according to the numbering. This can avoid the situation where the server of the target webpage receives a large number of requests for a long time, thus triggering the anti-crawling mechanism. This makes the number of requests received by the server of the target webpage change in the order of small, medium and large, improving the stability and intelligence of data crawling. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the system of the present invention;
[0051] Figure 2 This is a flowchart illustrating the steps of the method of the present invention;
[0052] Figure 3 This is a flowchart of the frequency testing loop of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1, First Aspect, Please refer to Figure 1 as well as Figure 3As shown, the present invention provides a web page data crawling system based on multiple channels, including a threshold testing module, a node grouping module, and a concurrent crawling module; the threshold testing module is used to obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output test nodes and crawling nodes;
[0055] It is also used to set different crawling frequencies for test nodes, crawl data from target websites, and output multiple crawling results; analyze multiple crawling results, and output frequency thresholds based on the analysis results;
[0056] The threshold testing module is configured with a node partitioning strategy, which includes:
[0057] Get the total number of proxy nodes and mark it as the total number of nodes;
[0058] Get the maximum crawling frequency of each proxy node and mark it as a node at full frequency;
[0059] Sort the proxy nodes according to their full frequency to obtain the node sequence;
[0060] Each proxy node is numbered in ascending order from left to right. Proxy nodes with numbers greater than the total number of nodes with the first scaling factor are marked as test nodes, and proxy nodes with numbers less than or equal to the total number of nodes with the first scaling factor are marked as crawling nodes.
[0061] In practice, the first scaling factor is set to 0.1. The purpose of setting the first scaling factor is to identify the proxy nodes with the highest crawling rate, which is used to test the maximum crawling frequency allowed by the webpage.
[0062] The threshold testing module is also configured with threshold testing strategies, which include:
[0063] Set the smallest number in the test node to i, and set n to an initial value of 0;
[0064] Execute the frequency test loop, which includes:
[0065] The crawling frequency of the test node with the number i is set to h×(n+1), where h is the first frequency. The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, no processing is performed; when the crawling fails, a failure message is output.
[0066] It should be noted that the crawler program should follow the target website's robots.txt file, which contains rules on what the website allows and does not allow crawling;
[0067] In practice, the first frequency is set to one-tenth of the full frequency of the test node with the highest number; taking Table 1-1 (number-frequency table) as an example:
[0068] Table 1-1 Number-Frequency Table
[0069] Node full frequency 50 60 70 80 90
[0070] Since the maximum allowed crawling frequency of the webpage to be crawled is unknown, let's assume it's 20. If we directly use test node number 5 at a crawling frequency of 50, test node number 5 will directly trigger the anti-crawling mechanism of the target webpage, resulting in a waste of IP resources. Subsequent adjustments at this point still have a high degree of randomness. Therefore, we set the first frequency to a smaller number. A preferred setting method is: one-tenth of the full frequency of the test node with the largest number, i.e., 9.
[0071] Using test node number 5 at a crawling frequency of 9, the frequency is safe and will not trigger anti-crawling mechanisms. Then, using test node number 6 at a crawling frequency of 18, it is still safe. Next, using test node number 7 at a frequency of 27, the first IP address used for testing is restricted. Finally, using test node number 6 at a frequency of 23, the second IP address used for testing is restricted, and 18 is set as the frequency threshold for the target webpage.
[0072] By using the threshold testing strategy, usually only one or two test nodes will trigger the anti-crawling mechanism and be restricted by IP, resulting in a waste of IP resources. At this time, the frequency threshold obtained is closer to the maximum crawling frequency allowed by the target webpage, thereby improving the crawling efficiency in the subsequent crawling process while avoiding the waste of IP resources.
[0073] If no failure message is received, increment n by one and execute the frequency test loop again.
[0074] When a failure message is received, n is decremented by one, and the crawling frequency of the test node numbered i+n is set to h×(n+1.5). The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, the crawling frequency is marked as the frequency threshold; when the crawling fails, h×(n+1) is set as the frequency threshold.
[0075] End the frequency test loop;
[0076] Mark the test nodes with numbers less than i+n back as crawling nodes.
[0077] The node grouping module is used to group crawled nodes, output multiple node groups, and number the node groups.
[0078] The node grouping module is configured with an interval grouping strategy, which includes:
[0079] The grouping process loop is executed, and the grouping process loop includes:
[0080] Set m to an initial value of 1;
[0081] The first random number of crawling nodes are assigned to the same group, which is marked as a node group and numbered m.
[0082] The second random number of crawling nodes are assigned to the same group, and the node group number is m×2;
[0083] The third random number of crawling nodes are assigned to the same group, and the node group is numbered m×3;
[0084] Determine if there are any remaining crawlable nodes. If the result is yes, perform a quantity check; if the result is no, end the group processing loop.
[0085] The quantity judgment process includes: judging whether the number of remaining crawling nodes is greater than the third number that is a multiple of the second. If the judgment result is yes, m is incremented by one, and the grouping process loop is executed again. If the judgment result is no, the remaining crawling nodes are divided into the same group, and the node group number is (m×3)+1.
[0086] In practical implementation, the second multiple means 2 times; taking a total of 40 crawled nodes as an example, the first number is set to 4, the second number is set to 8, and the third number is set to 12; the first to third numbers are set based on the number of proxy nodes in the actual application, and the specific settings should make the second number greater than or equal to twice the first number, and the third number greater than or equal to three times the first number.
[0087] It should be noted that using the same user agent every time data is scraped will trigger a warning signal that it is a bot; therefore, all crawling nodes should be grouped and controlled to crawl the target website in turn to avoid being identified as a bot.
[0088] In addition, to avoid the target webpage's server receiving a large number of requests for an extended period, which could overload the website server and lead to website crashes in extreme cases, all proxy nodes are divided into multiple node groups of varying numbers. The target webpage is crawled sequentially according to the node group number, allowing the server load of the target webpage to cycle through low, medium, and high load conditions.
[0089] The concurrent crawling module is used to set the crawling frequency of each crawling node based on a frequency threshold, control the concurrent crawling of crawling nodes in a node group based on the node group number, output crawling information or reassemble node information, and perform regrouping processing based on the reassembled node information; it is also used to integrate crawling information and output page information.
[0090] The concurrent crawling module is configured with concurrent crawling strategies, which include:
[0091] Set the crawling frequency of all crawled nodes to the frequency threshold of the second scaling factor;
[0092] In practice, the second scaling factor is set to 0.8. The purpose is that when the frequency threshold is exactly equal to the maximum allowed crawling frequency of the target webpage, all crawling nodes can crawl at a frequency less than the frequency threshold, thus reducing the risk of being identified.
[0093] Execute a concurrent crawling loop, which includes:
[0094] Reset m to 1 and n to 0. Control all proxy nodes in the node group with node number m to simultaneously crawl web page data using the crawler program according to the crawling frequency. When the crawl is successful, output the crawling information and assign the crawling information number n. Increment n and execute the completion judgment. When the crawl fails, end the concurrent crawling loop and output the recombined node information.
[0095] It should be noted that for some target web pages that limit the number of requests within a specific time period, such as 100 requests per minute, if the first number is set to 4, the second number to 8, the third number to 12, and the crawling frequency to 16, then the first group of nodes used for crawling will have 64 requests per minute, which is safe; however, the second group of nodes will have 128 requests per minute, exceeding the target web page's limit. In this case, it is necessary to further reduce the number of crawling nodes in each node group and output recombined node information.
[0096] The completion of the judgment process includes: determining whether the webpage data has been completely crawled; when the judgment result is yes, outputting crawling completion information and ending the concurrent crawling loop; when the judgment result is no, incrementing m by one and re-executing the concurrent crawling loop.
[0097] When reorganization node information is received, delete all node groups and scale the first, second, and third quantities by the second scaling factor.
[0098] Repeat the group processing loop and the concurrent crawling loop.
[0099] The concurrent crawling module is also configured with a page restoration strategy, which includes:
[0100] Sort all crawled information in ascending order of their numbers, mark all sorted crawled information as page information, and output the page information.
[0101] Example 2, Second Aspect, Please refer to Figure 2 As shown, the present invention also provides a multi-channel web page data crawling method, including:
[0102] Step S1: Obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output the test nodes and crawling nodes; Step S1 also includes the following sub-steps:
[0103] Step S101: Obtain the total number of proxy nodes and mark it as the total number of nodes;
[0104] Step S102: Obtain the maximum crawling frequency of each proxy node and mark it as a node at full frequency;
[0105] Step S103: Sort the proxy nodes according to their full frequency to obtain a node sequence;
[0106] Step S104: Number each proxy node in ascending order from left to right. Proxy nodes with numbers greater than the total number of nodes with the first scaling factor are marked as test nodes, and proxy nodes with numbers less than or equal to the total number of nodes with the first scaling factor are marked as crawling nodes.
[0107] Step S2 involves setting different crawling frequencies for the test nodes, crawling data from the target website, and outputting multiple crawling results; analyzing the multiple crawling results and outputting a frequency threshold based on the analysis results; Step S2 also includes the following sub-steps:
[0108] Step S201: Set the smallest number in the test node to i, and set n to an initial value of 0;
[0109] Step S202, execute the frequency test loop, which includes:
[0110] The crawling frequency of the test node numbered i+n is set to h×(n+1), where h is the first frequency. The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, no processing is performed; when the crawling fails, a failure message is output.
[0111] If no failure message is received, increment n by one and execute the frequency test loop again.
[0112] When a failure message is received, n is decremented by one, and the crawling frequency of the test node numbered i+n is set to h×(n+1.5). The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, the crawling frequency is marked as the frequency threshold; when the crawling fails, h×(n+1) is set as the frequency threshold.
[0113] End the frequency test loop;
[0114] Step S203: Mark the test nodes with numbers less than i+n back to the crawling nodes.
[0115] Step S3 involves grouping the crawling nodes, outputting multiple node groups, and numbering the node groups. A crawling frequency for each crawling node is set based on a frequency threshold. Concurrent crawling of the crawling nodes within a node group is controlled based on the group number. Crawling information or reconstructed node information is output, and regrouping is performed based on the reconstructed node information. Step S3 also includes the following sub-steps:
[0116] Step S301, execute the group processing loop, which includes:
[0117] Set m to an initial value of 1;
[0118] The first random number of crawling nodes are assigned to the same group, which is marked as a node group and numbered m.
[0119] The second random number of crawling nodes are assigned to the same group, and the node group number is m×2;
[0120] The third random number of crawling nodes are assigned to the same group, and the node group is numbered m×3;
[0121] Determine if there are any remaining crawlable nodes. If the result is yes, perform a quantity check; if the result is no, end the group processing loop.
[0122] The quantity judgment process includes: judging whether the number of remaining crawling nodes is greater than the third number that is a multiple of the second. If the judgment result is yes, increment m by one and execute the grouping process loop again; if the judgment result is no, divide the remaining crawling nodes into the same group and number the node group as (m×3)+1.
[0123] Step S302: Set the crawling frequency of all crawled nodes to the frequency threshold of the second scaling factor;
[0124] Step S303: Execute the concurrent crawling loop, which includes:
[0125] Reset m to 1 and n to 0. Control all proxy nodes in the node group with node number m to simultaneously crawl web page data using the crawler program according to the crawling frequency. When the crawl is successful, output the crawling information and assign the crawling information number n. Increment n and execute the completion judgment. When the crawl fails, end the concurrent crawling loop and output the recombined node information.
[0126] The completion of the judgment process includes: determining whether the webpage data has been completely crawled; when the judgment result is yes, outputting crawling completion information and ending the concurrent crawling loop; when the judgment result is no, incrementing m by one and re-executing the concurrent crawling loop.
[0127] Step S304: When the reorganization node information is received, delete all node groups and scale the first quantity, second quantity, and third quantity according to the second scaling factor.
[0128] Step S305: Execute the group processing loop and concurrent crawling loop again.
[0129] Step S4 involves integrating the crawled information and outputting the page information; Step S4 also includes the following sub-steps:
[0130] Step S401: Sort all crawled information in ascending order of number, mark all sorted crawled information as page information, and output the page information.
[0131] Example 3, a third aspect: This application provides an electronic device including a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the above method are performed. Through the above technical solution, the processor and memory are interconnected and communicate with each other via a communication bus and / or other forms of connection mechanism. The memory stores a computer program executable by the processor. When the electronic device is running, the processor executes the computer program to perform the method in any optional implementation of the above embodiments, to achieve the following functions: obtaining the total number of proxy nodes, dividing the proxy nodes, and outputting test nodes and crawling nodes; setting the crawling frequency, crawling data from the target website, and outputting the crawling results; analyzing the crawling results and outputting a frequency threshold; grouping the crawling nodes, setting the crawling frequency of each crawling node, controlling the crawling nodes to crawl concurrently, outputting crawling information or reorganizing node information, performing regrouping processing or integrating the crawling information, and outputting page information.
[0132] In embodiment 4, a fourth aspect, this application provides a storage medium storing a computer program thereon. When the computer program is executed by a processor, it performs the steps in the above method. Through the above technical solution, when the computer program is executed by a processor, it performs the method in any optional implementation of the above embodiments to achieve the following functions: obtaining the total number of proxy nodes, dividing the proxy nodes, and outputting test nodes and crawling nodes; setting the crawling frequency, crawling data from the target website, and outputting the crawling results; analyzing the crawling results and outputting a frequency threshold; grouping the crawling nodes, setting the crawling frequency of each crawling node, controlling the crawling nodes to crawl concurrently, outputting crawling information or reassembling node information, performing regrouping processing or integrating the crawling information, and outputting page information.
[0133] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0134] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0135] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
Claims
1. A multi-channel web page data crawling method, characterized in that, include: Get the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output the test nodes and crawl nodes; Set different crawling frequencies for test nodes, crawl data from the target website, and output multiple crawling results; analyze the multiple crawling results and output frequency thresholds based on the analysis results; The crawled nodes are grouped, and multiple node groups are output. The node groups are numbered. The crawling frequency of each crawling node is set based on the frequency threshold, and the crawling nodes in the node group are controlled to crawl concurrently based on the node group number. Crawling information or reconstructed node information is output, and regrouping is performed based on the reconstructed node information. The crawled information is integrated and the page information is output. Set different crawling frequencies for the test nodes, crawl data from the target website, and output multiple crawling results; Analyze multiple crawling results and output frequency thresholds based on the analysis results, including: Set the smallest number in the test node to i, and set n to an initial value of 0; Execute a frequency test loop, the frequency test loop including: Set the crawling frequency of the test node numbered i+n to h×(n+1), where h is the first frequency. Control the crawler program in the proxy node to crawl web page data according to the crawling frequency. When the crawling is successful, no processing is performed; when the crawling fails, a failure message is output. If no failure message is received, increment n by one and execute the frequency test loop again. When a failure message is received, n is decremented by one, and the crawling frequency of the test node numbered i+n is set to h×(n+1.5). The crawler program in the proxy node is controlled to crawl web page data according to the crawling frequency. When the crawling is successful, the crawling frequency is marked as the frequency threshold; when the crawling fails, h×(n+1) is set as the frequency threshold. End the frequency test loop; Mark the test nodes with numbers less than i+n back to the crawling nodes; The crawled nodes are grouped, and multiple node groups are output. The node groups are numbered as follows: Execute a group processing loop, the group processing loop including: Set m to an initial value of 1; The first random number of crawling nodes are assigned to the same group, which is marked as a node group and numbered m. The second random number of crawling nodes are assigned to the same group, and the node group number is m×2; The third random number of crawling nodes are assigned to the same group, and the node group is numbered m×3; Determine if there are any remaining crawlable nodes. If the result is yes, perform a quantity check; if the result is no, end the group processing loop. The quantity judgment process includes: judging whether the number of remaining crawling nodes is greater than the third number of the second multiple; if the judgment result is yes, increment m by one and execute the grouping processing loop again; if the judgment result is no, divide the remaining crawling nodes into the same group and number the node group as (m×3)+1.
2. The web page data crawling method based on multi-channel according to claim 1, characterized in that, Obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output the test nodes and crawling nodes, including: Get the total number of proxy nodes and mark it as the total number of nodes; Get the maximum crawling frequency of each proxy node and mark it as a node at full frequency; Sort the proxy nodes according to their full frequency to obtain the node sequence; Each proxy node is numbered in ascending order from left to right. Proxy nodes with numbers greater than the total number of nodes with the first scaling factor are marked as test nodes, and proxy nodes with numbers less than or equal to the total number of nodes with the first scaling factor are marked as crawling nodes.
3. The web page data crawling method based on multi-channel according to claim 2, characterized in that, The crawling frequency of each crawling node is set based on a frequency threshold, and the concurrent crawling of crawling nodes within a node group is controlled based on the node group number, including: Set the crawling frequency of all crawled nodes to the frequency threshold of the second scaling factor; Execute a concurrent crawling loop, the concurrent crawling loop including: Reset m to 1 and n to 0. Control all proxy nodes in the node group with node number m to simultaneously crawl web page data using the crawler program according to the crawling frequency. When the crawl is successful, output the crawling information and assign the crawling information number n. Increment n and execute the completion judgment. When the crawl fails, end the concurrent crawling loop and output the recombined node information. The completion judgment process includes: judging whether the web page data has been completely crawled; when the judgment result is yes, outputting crawling completion information and ending the concurrent crawling loop; when the judgment result is no, incrementing m by one and executing the concurrent crawling loop again.
4. The web page data crawling method based on multi-channel according to claim 3, characterized in that, Monitoring the crawling results and performing regrouping based on the monitoring results includes: When reorganization node information is received, delete all node groups and scale the first, second, and third quantities by the second scaling factor. Repeat the group processing loop and the concurrent crawling loop.
5. The web page data crawling method based on multi-channel according to claim 4, characterized in that, The crawled information is integrated, and the output page information includes: Sort all crawled information in ascending order of their numbers, mark all sorted crawled information as page information, and output the page information.
6. A multi-channel web page data crawling system, applicable to the multi-channel web page data crawling method described in any one of claims 1-5, characterized in that, It includes a threshold testing module, a node grouping module, and a concurrent crawling module; the threshold testing module is used to obtain the total number of proxy nodes, divide the proxy nodes based on the total number of proxy nodes, and output test nodes and crawling nodes. It is also used to set different crawling frequencies for test nodes, crawl data from target websites, and output multiple crawling results; analyze multiple crawling results, and output frequency thresholds based on the analysis results; The node grouping module is used to group crawled nodes, output multiple node groups, and number the node groups. The concurrent crawling module is used to set the crawling frequency of each crawling node based on a frequency threshold, control the concurrent crawling of crawling nodes in the node group based on the node group number, output crawling information or reassemble node information, and perform regrouping processing based on the reassembled node information. It is also used to integrate crawled information and output page information.
7. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-readable instructions that, when executed by the processor, perform the steps of the method as described in any one of claims 1-5.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it performs the steps of the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Data crawling method and system
CN116186368A
Webpage data crawling method, device and equipment and medium
CN109948026A
Webpage crawling method and device, computing equipment and storage medium
CN114610975A