Method for realizing data capture in combination with browser plug-in
By combining browser plug-in and Tampermonkey plug-in, VBS scripts are used to achieve automatic crawling and uploading of target website data, solving the problem of poor data collection stability caused by different website anti-crawling strategies, and improving the efficiency and quality of data crawling.
Patent Information
- Application Number
- CN202510089930.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
When facing anti-crawling strategies for different fields and types of websites, the existing technology requires a lot of time to analyze and crack, resulting in poor stability and maintenance of data collection and high learning costs.
By combining browser plug-ins, especially Chrome plug-in extension methods, Tampermonkey plug-in and VBS scripts can be used to automatically crawl and upload the target website data, and avoid anti-crawling strategies.
Effectively bypass the anti-crawl strategy, improve the efficiency and quality of data crawling, reduce learning costs and operational complexity, realize automated regular crawling tasks, and reduce the risk of being identified.
Smart Images

Figure CN120011620A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of data collection, and specifically is a method for realizing data capture by combining with a browser plug-in. Background Art
[0002] In the era of big data, collecting information is an important task. If information is collected solely by manpower, it will not only be inefficient and cumbersome, but the cost of collection will also increase.
[0003] Web crawlers, also called network robots, can automatically collect and organize data information on the Internet instead of people. Staff can use web crawlers to collect data of interest on specific target websites, and apply them to multiple application scenarios such as data analysis and public opinion monitoring. However, different types of websites in different fields will have different anti-crawling strategies.
[0004] The first is client-side anti-crawl. The front-end encrypts the parameters of the requested API and loads it dynamically. The website will use related front-end anti-crawl technologies such as ajax dynamic loading of content, browser camouflage identification technology, anti-debugging, code obfuscation, etc. The second is server-side anti-crawl. The website will identify and verify the access frequency, access characteristics, access IP, etc., and perform secondary verification of human-machine identification based on the verification code through big data analysis, artificial intelligence and other technologies.
[0005] At present, in order to crawl the encrypted data of the website, it is usually necessary to analyze the preconditions for collecting data on a specific website, and then analyze the anti-crawling strategy of the website and crack them one by one, so that the desired data can be finally obtained. However, due to the different fields and types of websites, their anti-crawling strategies are different. It takes a lot of time to analyze and crack their respective anti-crawling methods, and the learning cost is high. If the data structure and anti-crawling strategy change, it will lead to poor stability and maintainability of data collection. Therefore, a solution that can stably realize data crawling is needed for optimization. Summary of the invention
[0006] The purpose of the present invention is to provide a method for implementing data crawling in combination with a browser plug-in to solve the problems raised in the above background technology.
[0007] In order to achieve the above object, the present invention provides the following technical solution: a method for realizing data crawling in combination with a browser plug-in, which is based on the Chrome plug-in extension and realizes crawling of target website data by simply configuring the address of the website to be crawled and the data interface URL, and the specific steps of the solution are:
[0008] Step 1: Task plan configuration: configure the scheduled task on the deployment server;
[0009] Step 2: Task information configuration: extract and configure the server data of the target website;
[0010] Step 3: The crawling task is executed, the browser is started through the VBS script, and the website page is jumped;
[0011] Step 4: Capture and save data, match and capture the required data information for saving;
[0012] Step 5: Data upload: upload the collected data to the remote web server.
[0013] Preferably, during the task scheduling execution phase, a Windows scheduled task is used to periodically execute a VBS script, and the VBS script is responsible for starting a browser to perform a crawling task.
[0014] Preferably, the task information configuration stage configures the server data into the Chrome plug-in, wherein the server data includes the target address to be captured, the data interface address of the target, and the server address information where the data is stored.
[0015] Preferably, the specific steps of the crawling task execution phase are:
[0016] A1, Windows scheduled task starts and automatically executes the VBS script;
[0017] A2, the VBS script will start the browser, and when the browser is started, the Chrome plug-in will execute the Tampermonkey plug-in through the configured target website address;
[0018] A3, Tampermonkey plug-in will redirect the page according to the target website address and monitor all network requests;
[0019] A4, at this time, you can use the Chrome plug-in to capture data information.
[0020] Preferably, the VBS script can open a real Chrome browser, and capture data by jumping to a page in combination with a Tampermonkey plug-in, thereby avoiding the website's blocking and restriction of simulated browser collection.
[0021] Preferably, during the data capture and saving stage, the Chrome plug-in compares the address information in the address bar with the data interface address in the configuration file, and when the addresses match, captures the required information for saving.
[0022] Preferably, in the data uploading stage, the saved data information is asynchronously submitted to the remote web server via http Post, completing the entire process of data capture.
[0023] Preferably, the Windows scheduled task can set the execution interval of the VBS script and is responsible for simulating the normal access frequency.
[0024] The beneficial effects of the present invention are as follows:
[0025] 1. The present invention starts a real Chrome browser and combines it with the Tampermonkey plug-in to crawl, thereby avoiding the interception of common anti-crawling mechanisms such as simulated browser blocking and IP restriction, and can effectively bypass the anti-crawling strategy, thereby improving the efficiency of data crawling. Compared with the traditional method, the present invention only needs to configure various data of the server through the configuration file to specify and quickly collect the target website data. More importantly, the entire crawling process can be quickly configured and adapted to different websites and data sources, the operation is simple, and the learning cost is reduced.
[0026] 2. The present invention executes VBS scripts through Windows scheduled tasks, so that after the data crawling task is configured, the crawling task can be automatically and regularly executed without manually starting the script, thereby reducing the crawling burden. More importantly, when crawling for a long time, the execution interval of the VBS script can be set through the Windows scheduled task to simulate the normal access frequency and avoid triggering the secondary verification of the anti-crawling mechanism, thereby ensuring the smoothness of data crawling, reducing the risk of being identified, and making the crawling efficiency high and the interference less.
[0027] 3. The present invention compares the URL in the browser address bar and the data interface address in the configuration file through the Tampermonkey plug-in to ensure that the captured data conforms to the predetermined rules and can accurately match the data to avoid capturing irrelevant data, thereby improving the quality and accuracy of data capture. The data is then asynchronously uploaded to the remote server through http Post, ensuring the efficiency of data upload without blocking the capture task, thereby improving applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a data capture flow chart of the present invention;
[0029] Figure 2 This is a flow chart of the operation of the system of the present invention;
[0030] Figure 3 This is an example diagram of the source code of the web page being crawled according to the present invention. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] like Figures 1 to 3 As shown, an embodiment of the present invention provides a method for realizing data crawling in combination with a browser plug-in. The solution is based on a Chrome plug-in extension and realizes crawling of target website data by simply configuring the address of the website to be crawled and the data interface URL. The specific steps of the solution are:
[0033] Step 1: Task plan configuration: configure the scheduled task on the deployment server;
[0034] Step 2: Task information configuration: extract and configure the server data of the target website;
[0035] Step 3: The crawling task is executed, the browser is started through the VBS script, and the website page is jumped;
[0036] Step 4: Capture and save data, match and capture the required data information for saving;
[0037] Step 5: Data upload: upload the collected data to the remote web server.
[0038] By launching a real Chrome browser and combining it with the Tampermonkey plug-in for crawling, we avoid the interception of common anti-crawling mechanisms such as simulated browser blocking and IP restrictions, and can effectively bypass anti-crawling strategies and improve the efficiency of data crawling. Compared with traditional methods, this solution only needs to configure various server data through configuration files to specify and quickly collect target website data. More importantly, the entire crawling process can be quickly configured and adapted to different websites and data sources. The operation is simple and the learning cost is reduced.
[0039] In the task scheduling execution stage, a Windows scheduled task is used to periodically execute a VBS script, and the VBS script is responsible for starting a browser to perform a crawling task.
[0040] During task scheduling, the VBS script is responsible for starting the browser and triggering subsequent crawling tasks. Through the VBS script, the browser startup process can be controlled to ensure the execution order and stability of the crawler tasks.
[0041] Among them, the task information configuration stage will configure the server data into the Chrome plug-in, where the server data includes the target address to be captured, the target data interface address, and the server address (URL address) information where the data is saved.
[0042] By configuring server data, you can modify the configuration file according to different crawling tasks without modifying the core script to achieve rapid customization.
[0043] The specific steps of the crawling task execution phase are as follows:
[0044] A1, Windows scheduled task starts and automatically executes the VBS script;
[0045] A2, the VBS script will start the browser, and when the browser is started, the Chrome plug-in will execute the Tampermonkey plug-in through the configured target website address;
[0046] A3, Tampermonkey plug-in will redirect the page according to the target website address and monitor all network requests;
[0047] A4, at this time, you can use the Chrome plug-in to capture data information.
[0048] The VBS script is responsible for automatically starting the real Chrome browser and triggering the Tampermonkey plug-in. At this time, the Tampermonkey plug-in executes page jumps, loads related data, and dynamically obtains data according to the target address in the configuration file.
[0049] By executing VBS scripts through Windows scheduled tasks, data crawling tasks can be automatically executed regularly after configuration is completed, without the need to manually start the script, thus reducing the crawling burden. More importantly, when crawling for a long time, the execution interval of the VBS script can be set through Windows scheduled tasks to simulate the normal access frequency and avoid triggering the secondary verification of the anti-crawling mechanism, thus ensuring the smoothness of data crawling, reducing the risk of being identified, and making crawling efficient with less interference.
[0050] The VBS script can open a real Chrome browser and use the Tampermonkey plug-in to jump to the page to capture data, thereby avoiding the website's blocking and restrictions on simulated browser collection.
[0051] Many modern websites rely on JavaScript to dynamically load data, and simulated crawlers cannot directly handle such requests. By running in a real browser environment, the risk of blocking websites by detecting the simulated behavior of crawler tools can be avoided. Compared with common headless browser simulation, this method can simulate user behavior more realistically and reduce the recognition of anti-crawling mechanisms.
[0052] Among them, in the data capture and saving stage, the Chrome plug-in will compare the address information in the address bar with the data interface address in the configuration file, and when the addresses match, the required information will be captured and saved.
[0053] Among them, in the data uploading stage, the saved data information is asynchronously submitted to the remote web server through the http Post method to complete the entire process of data capture.
[0054] By using asynchronous uploading, the crawling task will not be blocked when the data is being uploaded, so large amounts of data can be processed efficiently.
[0055] The Tampermonkey plug-in is used to compare the URL in the browser address bar with the data interface address in the configuration file to ensure that the captured data conforms to the predetermined rules and can accurately match the data to avoid capturing irrelevant data, thereby improving the quality and accuracy of data capture. The data is then asynchronously uploaded to the remote server through http Post, ensuring the efficiency of data upload without blocking the capture task and improving applicability.
[0056] The Windows scheduled task can set the execution interval of the VBS script and is responsible for simulating the normal access frequency.
[0057] When the access frequency is very high, the anti-crawling mechanism of some servers will perform a second verification. The Windows scheduled task can ensure the normal progress of the data crawling task.
[0058] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0059] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for realizing data capture in combination with a browser plug-in, characterized in that: This solution is based on the Chrome plug-in extension and achieves crawling of target website data by simply configuring the address of the website to be crawled and the data interface URL. The specific steps of this solution are: Step 1: Task plan configuration: configure the scheduled task on the deployment server; Step 2: Task information configuration: extract and configure the server data of the target website; Step 3: The crawling task is executed, the browser is started through the VBS script, and the website page is jumped; Step 4: Capture and save data, match and capture the required data information for saving; Step 5: Data upload: upload the collected data to the remote web server.
2. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: During the task scheduling execution phase, Windows scheduled tasks are used to periodically execute VBS scripts, and the VBS scripts are responsible for starting the browser to perform crawling tasks.
3. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: The task information configuration stage configures the server data into the Chrome plug-in, where the server data includes the target address to be captured, the target data interface address, and the server address information where the data is saved.
4. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: The specific steps of the crawling task execution phase are: A1, Windows scheduled task starts and automatically executes the VBS script; A2, the VBS script will start the browser, and when the browser is started, the Chrome plug-in will execute the Tampermonkey plug-in through the configured target website address; A3, Tampermonkey plug-in will redirect the page according to the target website address and monitor all network requests; A4, at this time, you can use the Chrome plug-in to capture data information.
5. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: The VBS script can open a real Chrome browser and use the Tampermonkey plug-in to jump to the page to capture data, thereby avoiding the website's blocking and restrictions on simulated browser collection.
6. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: During the data capture and saving stage, the Chrome plug-in compares the address information in the address bar with the data interface address in the configuration file, and when the addresses match, captures the required information for saving.
7. The method for realizing data capture by combining with a browser plug-in according to claim 1, characterized in that: In the data uploading stage, the saved data information is asynchronously submitted to the remote web server via http Post, completing the entire process of data capture.
8. The method for realizing data capture by combining with a browser plug-in according to claim 2, characterized in that: The Windows scheduled task can set the execution interval of the VBS script and is responsible for simulating the normal access frequency.
Citation Information
Cited By
System for automatically acquiring page table data based on process
CN120950778A