System for automatically acquiring page table data based on process

By using a structure recognition plugin and a user-visual configuration interface, web page table data can be automatically captured, solving the problems of frequent rule failures and high maintenance costs in traditional methods, and achieving efficient and stable data acquisition.

CN120950778APending Publication Date: 2025-11-14SHANDONG INSPUR SCI RES INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511206800.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional web scraping methods rely on manually writing parsing rules, which leads to frequent rule failures, high maintenance costs, high technical requirements for users, and a lack of automated recognition capabilities, making it difficult to adapt to the complexity and variability of modern web page structures.

Method used

This invention provides a system for automating the acquisition of page table data, including a structure recognition plugin, a user-visual configuration interface, and a data capture module. By automatically recognizing the webpage structure, configuring data collection rules, and executing data capture and storage, it lowers the technical threshold for users and adapts to changes in webpage structure.

Benefits of technology

It enables automated scraping of table data from web pages, improving data acquisition efficiency and stability, reducing maintenance costs, adapting to changes in web page structure, and eliminating the need for users to write complex parsing rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950778A_ABST
    Figure CN120950778A_ABST
Patent Text Reader

Abstract

The invention provides a system for automatically acquiring page table data based on a process, which belongs to the technical field of distributed file transmission and integrally comprises a structure identification plug-in module, a user visual configuration interface module, a data capture engine module and a data post-processing module. A webpage table structure is recognized through a plug-in, a user configures rules in a graphical interface, and grabbing and storage are automatically executed. According to the system, on the premise that complex rules do not need to be written manually, webpage data collection automation and high availability are achieved, the data obtaining efficiency is improved by about 80% compared with a traditional mode, and the maintenance cost is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed file transfer technology, and more particularly to a system for automating the acquisition of page table data, which is especially suitable for scenarios requiring efficient transfer of large-scale files, including software version updates, container image distribution, multimedia content transmission, and firmware upgrades for IoT devices. This invention significantly improves file distribution efficiency through an innovative hierarchical storage architecture and dynamic routing optimization mechanism. Background Technology

[0002] Traditional web scraping methods primarily rely on manually writing parsing rules. This approach requires users to have a deep understanding of the webpage's HTML structure and to write corresponding code to extract the desired data based on the specific content of the webpage. However, modern webpages are typically very complex, containing a large amount of dynamic content, nested structures, and various data formats. Manually writing parsing rules is not only time-consuming and laborious but also struggles to cope with frequent changes in webpage structure.

[0003] This leads to the following problems:

[0004] Frequent rule invalidation: Once the webpage structure changes (e.g., webpage layout adjustment, HTML tag change, data loading method change, etc.), manually written parsing rules will become invalid, resulting in data crawling failure or errors.

[0005] High maintenance costs: To maintain the stability of data crawling, parsing rules need to be constantly updated and maintained. This not only requires specialized technical knowledge but also a significant investment of manpower and time, increasing the overall cost of data crawling.

[0006] The reasons for this can be divided into the following aspects:

[0007] Web page structures are complex and constantly changing: Modern web pages typically employ dynamic loading technologies (such as AJAX, JavaScript dynamic rendering, etc.), and their HTML structure may change at different points in time. Furthermore, the design and layout of web pages may change due to version updates, content adjustments, and other reasons, making manually written parsing rules difficult to adapt to these changes.

[0008] Lack of automated recognition capabilities: Traditional methods cannot automatically recognize the structure and content of web page table data, requiring manual intervention to adjust the parsing rules.

[0009] High technical requirements for users: Manually writing parsing rules requires users to have certain programming skills and a deep understanding of HTML, CSS, JavaScript, and other technologies. This limits the use of web scraping tools by many non-technical users and reduces the accessibility of data acquisition. Summary of the Invention

[0010] To address the aforementioned technical challenges, this invention provides a system for automating the acquisition of webpage table data. With this system, users can automate the retrieval of webpage table data without writing complex parsing rules, requiring only simple configuration. This technology can automatically identify the structure and content of webpage tables, adapting to changes in webpage structure, thereby significantly improving the efficiency and stability of data acquisition and reducing maintenance costs.

[0011] The technical solution of this invention is:

[0012] A system for automating the acquisition of page table data based on a workflow, including

[0013] The structure recognition plugin automatically identifies the data table areas that can be collected on the webpage after the browser loads the webpage, and guides the user to configure data collection in a graphical form.

[0014] The user-visual configuration interface displays the results of webpage structure recognition and provides a WYSIWYG (What You See Is What You Get) approach to configuring data collection rules, enabling automatic mapping and manual verification from "webpage table structure" to "structured extraction parameters".

[0015] The data capture module is responsible for automatically executing the entire process of webpage loading, element location, pagination navigation, data extraction, data cleaning, and storage, based on the output of the first two modules, namely the structure recognition plugin and the user visual configuration interface.

[0016] The data post-processing module is responsible for persistently storing the structured data extracted in the previous stage, and performing subsequent format cleaning, structure adjustment, visualization processing, and multi-channel export to provide users or systems with directly usable data results, and supports integration and reuse with low-code platforms.

[0017] Furthermore,

[0018] The working steps of the structure recognition plugin are as follows:

[0019] Step 1: Plugin Injection and Page Load Listening

[0020] Develop browser extensions, define content_scripts configurations, and automatically inject scripts when a user visits a target website;

[0021] Use window.onload, DOMContentLoaded, or MutationObserver to listen for when the webpage has finished loading;

[0022] To prevent partial page refreshes from causing delayed data loading, a method is used to determine when the structure is ready to load by combining setInterval and DOM change listeners.

[0023] Step 2: DOM Structure Parsing and Candidate Region Extraction

[0024] Get all of the pages

[0025] 、 Tags;

[0026] Extract features from its child nodes, such as text density, row and column structure, and whether there is a title, and construct a list of candidate data regions.

[0027] Step 3: Data Structure Pattern Analysis and Field Identification

[0028] Iterate through each row of the candidate region and extract the text content of each column;

[0029] Compare whether the first row of content has the characteristics of a field title;

[0030] Analyze the data types of all columns and construct a field type matrix;

[0031] Step 4: Visual highlighting and numbering of candidate regions

[0032] Use front-end technologies to add borders and semi-transparent masks to the candidate table area;

[0033] Add a numbered label to the top left corner of each candidate area, which users can click to select;

[0034] When the mouse hovers over the area, a preview of the first few rows of data extracted from that area is displayed.

[0035] Step 5: User selection interaction and field confirmation guidance

[0036] After the user clicks on the highlighted area, a configuration panel will pop up;

[0037] The panel displays the table field names and data previews;

[0038] Users can adjust the field mapping relationships;

[0039] Supports operations such as renaming fields, enabling / disabling fields, and setting field types;

[0040] Step Six: Generate Structure Rules and Initialize Data Acquisition Tasks

[0041] Generation method:

[0042] Abstract the information from DOM selectors and field indexes into general rules;

[0043] If pagination is enabled, a corresponding pagination logic description will be generated.

[0044] Rules can be stored in the plugin's local storage, IndexedDB, or sent to the backend via API.

[0045] in,

[0046] Step two uses heuristic rules, including:

[0047] The average number of child nodes per row is ≥2;

[0048] The number of child nodes in all rows is basically the same;

[0049] There is one line with a clear title word;

[0050] Machine learning models can be introduced to predict the structure of historical structural samples.

[0051] Step 3 extracts the result structure into a two-dimensional array format:

[0052] Utilizes NLP for field name recognition, compatible with Chinese and multiple languages;

[0053] Use regular expressions to determine the type of specific fields such as date, amount, and phone number.

[0054] In step five, the configuration interface can embed a drawer bar on the right side of the page to avoid redirection;

[0055] All user configurations are saved in JSON rule format for easy data retrieval and use later.

[0056] Furthermore,

[0057] The user-visual configuration interface and the workflow are as follows:

[0058] Step 1: Initializing and displaying the rule configuration panel

[0059] After the user clicks on the identified table area or clicks the data collection configuration button, a configuration panel pops up on the side of the page;

[0060] The panel can be built using HTML+CSS+JavaScript / Vue / React and adopts a responsive design;

[0061] After the panel loads, it automatically fills in the data fields returned by the structure recognition module and displays a table preview.

[0062] Step 2: Field Parameter Mapping and Field Editing Interaction

[0063] Core functionalities:

[0064] Field name settings: Displays the title of each column, and supports renaming;

[0065] Field Enable / Disable: Control whether to collect data for this field via a checkbox;

[0066] Field type selection: Provides a drop-down list for users to select the field type;

[0067] Field Preview: This feature displays sample data for the field from the first few rows, helping to determine its meaning.

[0068] Step 3: Configure pagination rules

[0069] The configuration interface offers options for pagination methods;

[0070] The system automatically captures the DOM selector when a user clicks the "Next Page" button on the page.

[0071] You can also manually enter the CSS selector or XPath for the pagination button;

[0072] Set the maximum number of page turns or the termination condition;

[0073] Step 4: Data Preview and Rule Testing

[0074] According to the current rules, the first N rows of data are extracted and displayed in real time at the bottom of the configuration interface.

[0075] If pagination exists, you can manually test page turning and confirm the crawling effect;

[0076] Error messages can be displayed;

[0077] Step 5: Rule naming, saving, and reuse mechanism

[0078] The configuration interface provides a rule name input box at the top;

[0079] Clicking the save button will save the current rule structure to local storage, synchronize it with the browser, or upload it to the server.

[0080] It supports three operation modes: creating new rules, saving overwrite rules, and importing existing rules.

[0081] Step Six: Rule Format Validation and Task Submission Guidance

[0082] Implementation method:

[0083] When the "Submit Data Collection Task" button is clicked, the completeness of the rules is verified:

[0084] Are any valid tables selected?

[0085] Is at least one field configured?

[0086] Is the pagination setting reasonable?

[0087] After successful verification, a task confirmation window will be displayed to guide the user:

[0088] Naming tasks;

[0089] Select the execution method;

[0090] Submit the task to the scheduling engine.

[0091] in,

[0092] In step one, front-end frameworks such as Ant Design, Element UI, and Vuetify can be used to quickly build the UI; the panel should support drag and drop and automatically adapt to the current screen size.

[0093] In step two, use tooltips to provide field type descriptions; allow setting default values ​​or data cleaning rules;

[0094] In step four, all preview results are not written to the database, but are rendered only in memory; asynchronous data requests are used to avoid page lag; and preview results can be exported as JSON / CSV sample files.

[0095] In step five, save it as a frequently used data collection template for the user account; set rule tags for easy categorization and searching.

[0096] Furthermore,

[0097] The data scraping module works as follows:

[0098] Step 1: Task Scheduling and Rule Loading

[0099] The scheduling system inputs task parameters, including the target URL, collection rule ID, and collection time;

[0100] The module parses the JSON file containing the rules, reads the field location rules, pagination methods, and collection conditions; and initializes the log tracker to record the status of the crawling process.

[0101] Step Two: Webpage Loading and Environment Preparation

[0102] Use a headless browser to simulate real user browsing behavior;

[0103] Configure user agent and anti-scraping mechanism bypass strategies;

[0104] Wait for the target table element to finish loading;

[0105] Supports setting timeout and retry mechanisms to avoid interrupting the process due to loading failure;

[0106] Step 3: Execution of Master Data Extraction Logic

[0107] Traverse the table rows in the webpage ( Tags; simultaneously scanning for specific structural features Extract each column according to the rules. The fields required in the configuration; for each field, extract the text, attribute value, or link href according to the configured selector;

[0108] Supports field-level data processing: whitespace removal, unit conversion, regular expression extraction, and data standardization; automatic type validation for corresponding fields improves data quality.

[0109] Step 4: Pagination Processing and Merging of Multiple Pages

[0110] After retrieving data from the current page, find the pagination control element according to the rules;

[0111] Perform the page turning action, wait for the page to refresh, and detect changes in the table;

[0112] The process of extracting duplicate data continues until the maximum number of pages is reached or a termination condition is detected.

[0113] Step 5: Data Cleaning and Structure Preparation

[0114] Null value handling: Clear blank rows and empty fields;

[0115] Type conversion: Price converted to floating-point; Date format standardized;

[0116] Regular expression extraction: Extracting zip codes from addresses and brands from titles;

[0117] Duplicate data detection: Remove duplicates based on the primary key field;

[0118] Step Six: Data Storage and Export

[0119] Supported storage methods:

[0120] Store in the database;

[0121] Upload to the backend API interface;

[0122] Exporting local files;

[0123] Send to a third-party platform via middleware.

[0124] Furthermore,

[0125] The working steps of the data post-processing module are as follows:

[0126] Step 1: Structured Data Encapsulation

[0127] Encapsulate the data into a JSON or DataFrame structure, with each record having a unified field key-value pair; add metadata fields;

[0128] Type-mark the data items;

[0129] Construct a mapping relationship between standard model fields and original webpage fields;

[0130] Step 2: Data Cleaning and Standardization

[0131] Deduplication: Remove duplicate records based on the primary key field;

[0132] Standardized format;

[0133] Step 3: Persistent Data Storage

[0134] Call the database driver or ORM tool to perform batch inserts;

[0135] Automatic table and database creation, with automatic field type matching;

[0136] Creating indexes improves retrieval efficiency;

[0137] Data storage logs record operation results and exceptions;

[0138] Step 4: Exporting Results and Supporting Multiple Formats

[0139] Export logic:

[0140] Users select fields, time ranges, and export formats through the interface;

[0141] The module automatically converts the corresponding data into the target format;

[0142] For large data exports, use pagination or asynchronous compression and packaging.

[0143] Step 5: Building Reusable Interfaces for Low-Code Platforms

[0144] Provides a RESTful API interface that supports GET queries and POST submissions;

[0145] Supports GraphQL querying, enhancing front-end integration flexibility;

[0146] Configurable interface data structure, pagination parameters, and field filtering conditions;

[0147] Supports OAuth2 and token-based access control;

[0148] Provides a data source registration function, allowing you to add sources directly in the low-code platform;

[0149] Integrate a data change notification mechanism to enable automatic refresh;

[0150] Step Six: Data Collection Result Traceability and Data Security Assurance

[0151] Implementation mechanism:

[0152] Data entries are bound to task IDs, user IDs, and source URLs for easy traceability.

[0153] Each storage action is automatically logged, and rollback is supported.

[0154] Sensitive information is automatically encrypted or de-identified;

[0155] Configure backup strategies for the storage system to prevent data loss.

[0156] in,

[0157] In step three, the supported data storage types are: relational databases, NoSQL databases, object storage, and local files.

[0158] Step 4: Exporting Results and Supporting Multiple Formats

[0159] Export format support:

[0160] CSV / Excel is used in office scenarios;

[0161] JSON / XML is used to interface with system APIs;

[0162] Markdown / HTML tables are used for report display;

[0163] Template export is supported. Attached Figure Description

[0164] Figure 1 This is a schematic diagram of the system architecture of the present invention, showing the relationship between the plug-in, configuration interface, data capture module and data storage module;

[0165] Figure 2 This is a schematic diagram of the configuration interface, showing how users can specify the tables to be captured through the configuration interface;

[0166] Figure 3 This is a diagram illustrating the data scraping process, showing in detail how the plugin parses web pages and scrapes data based on user configuration. Detailed Implementation

[0167] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0168] This invention proposes a system for acquiring webpage table data based on process automation. It adopts an overall architecture of "plugin + configuration interface + automatic crawling engine + data storage and processing" to build a general-purpose data acquisition tool for structured webpage data (especially table data), helping users to automatically identify, quickly configure and automatically collect target data from webpages.

[0169] This invention is applicable to a wide range of scenarios, such as e-commerce data collection, public information capture, automatic entry into industry databases, and extraction of regulatory information. It is particularly valuable in industries such as government, finance, telecommunications, and power, where data accuracy, collection frequency, and changes are high and frequent.

[0170] 1. Plugin Development: Specific Implementation Steps for the Structure Recognition and Task Guidance Module.

[0171] This module serves as the "sentinel sensor" of the entire invention system. Its goal is to automatically identify the data table areas that can be collected on the webpage after the browser loads the webpage, and guide the user to configure data collection in a graphical form.

[0172] Step 1: Plugin Injection and Page Load Listening

[0173] Objective: To automatically start the plugin logic after the webpage has finished loading.

[0174] Implementation method:

[0175] Develop browser extensions (such as Chrome Extensions) and define content_scripts configurations to automatically inject scripts when a user visits a target website.

[0176] Use window.onload, DOMContentLoaded, or MutationObserver to listen for when a webpage finishes loading.

[0177] To prevent data loading delays caused by partial page refreshes, a structure loading readiness check is implemented by combining setInterval and DOM change listeners.

[0178] Code example:

[0179] document.addEventListener('DOMContentLoaded',()=>{

[0180] initTableRecognition();

[0181] });

[0182] Step 2: DOM Structure Parsing and Candidate Region Extraction

[0183] Objective: To filter out areas in the entire DOM structure that may be "data tables".

[0184] Implementation method:

[0185] Get all of the pages

[0186] 、 Tags such as "grid", "table", "list", etc. (e.g., classes containing "grid", "table", "list", etc.);

[0187] Extract features such as text density, row and column structure, and whether there is a title from its child nodes to construct a "candidate data region list".

[0188] Technical points:

[0189] The candidate area of ​​the table is not limited to standard HTML Tags; simultaneously scanning for specific structural features

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0203]

[0204]

[0205]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242]

[0243]

[0244]

[0245] The following steps need to be considered: "Pseudo-table structure"; use heuristic rules such as: average number of child nodes per row ≥ 2; the number of child nodes in all rows is basically the same; there exists a row with obvious title words (such as "name", "price", "time"); machine learning models can be introduced to predict the structure of historical structure samples (optional). Step 3: Data Structure Pattern Analysis and Field Identification Objective: To identify data fields and their structure (field names, data columns, etc.) in the candidate area. Implementation Method: Traverse each row (tr, div.row, etc.) of the candidate area and extract the text content of each column; compare whether the content of the first row has the "field title" characteristic, such as whether it is a Chinese noun or whether it is inconsistent with the next row; count the data types of all column contents (such as pure numbers, dates, strings) and construct a field type matrix; the extracted result structure is a two-dimensional array format: [["Name","Price","Inventory"],["Apple","3.50","100"],["Banana","2.80","80"]]. Optional Enhancements: Use NLP for field name recognition, compatible with Chinese and multiple languages; use regular expressions to determine specific field types such as date, amount, and phone number. Step Four: Visual Highlighting and Numbering of Candidate Areas Purpose: To visualize the selectable data table areas on the page for easy user configuration. Implementation Method: Use front-end technologies (such as JavaScript + CSS) to add borders and semi-transparent overlays to the candidate table areas; add number labels to the upper left corner of each candidate area, which users can click to select; display a preview of the first few rows of data extracted for that area when the mouse hovers over it. Code Example: element.style.border = "2px solid#4CAF50"; element.setAttribute("data-table-id", index); Interaction Design Notes: Highlighting should not affect the original functionality of the page; the highlighted area should accurately match the table boundary (boundingRect can be used to assist in the judgment); support "deselect" and "re-identify". Step 5: User Selection Interaction and Field Confirmation Guidance Purpose: Guide users to select the target table from multiple candidate areas and confirm the fields. Implementation Method: After the user clicks the highlighted area, a configuration panel pops up; the panel displays the table field names and data preview; users can adjust the field mapping relationship (such as "Column 1 → Product Name", "Column 3 → Unit Price"); supports field renaming, enabling / disabling fields, setting field types, etc. Key Technical Points: The configuration interface can be embedded in the right drawer bar of the page to avoid navigation; all user configurations are saved in JSON rule format for easy use in subsequent data crawling. Step 6: Generate Structure Rules and Initialize Collection Task Purpose: Generate structure rules that can be executed by the collection engine from the information finally selected and configured by the user.Example rule format (JSON): Generation method: Abstracts DOM selectors and field indexes into general rules; if pagination is enabled, generates corresponding pagination logic descriptions; rules can be stored in the plugin's local storage (LocalStorage), IndexedDB, or transmitted to the backend via API. Technical innovations: Strong dynamic structure recognition capability: No need to rely on fixed webpage templates, supports webpage tables with different structures; Pseudo-table support: Breakthrough. Data regions constructed using structures such as [list of structures].

[0246] User-friendly configuration experience: Through visual guidance and interactive configuration, data scraping rule design can be completed without writing code;

[0247] Scalable and reusable crawling rule mechanism: Structure rules can be reused, shared, or used as independent configuration files for scheduling platforms.

[0248] 2. Configuration Interface: The configuration interface is designed to be user-friendly, allowing users to specify the tables to be scraped through simple configuration.

[0249] This module is used to display the results of webpage structure recognition to users and provides a WYSIWYG (What You See Is What You Get) method for configuring data collection rules, enabling automatic mapping and manual confirmation from "webpage table structure" to "structured extraction parameters".

[0250] Step 1: Initializing and displaying the rule configuration panel

[0251] Objective: To provide a configuration entry point and pop up an interactive configuration interface.

[0252] Implementation method:

[0253] After the user clicks on the identified table area or clicks the "Collection Configuration" button, a configuration panel (sidebar / pop-up box) will pop up on the side of the page.

[0254] The panel can be built using HTML+CSS+JavaScript / Vue / React, and responsive design is recommended;

[0255] After the panel loads, it automatically fills in the data fields returned by the structure recognition module and displays a table preview.

[0256] Code example:

[0257]

[0258] Technical recommendations:

[0259] You can quickly build UIs using front-end frameworks such as Ant Design, Element UI, and Vuetify;

[0260] The panel should support dragging and automatically adapt to the current screen size.

[0261] Step 2: Field Parameter Mapping and Field Editing Interaction

[0262] Purpose: To allow users to modify or confirm field names, types, and enabled status.

[0263] Core functionalities:

[0264] Field name settings: Displays the title of each column, and supports renaming;

[0265] Field Enable / Disable: Control whether to collect data for this field via a checkbox;

[0266] Field type selection: Provides a drop-down list for users to select the field type (text, number, date, boolean, etc.);

[0267] Field Preview: This feature displays sample data for the field from the first few rows, helping to determine its meaning.

[0268]

[0269] {"fieldName":"image address","type":"url","enabled":false} ]

[0271] Enhanced design:

[0272] Use tooltips to provide descriptions of field types;

[0273] Allows setting default values ​​or data cleaning rules (such as removing the word "yuan").

[0274] Step 3: Configure pagination rules

[0275] Objective: To extract all page data from a website with pagination.

[0276] Supported pagination types:

[0277] Click the "Next Page" button;

[0278] Dropdown loading;

[0279] URL parameter pagination (e.g., ?page=2);

[0280] Page number selector within a table.

[0281] Implementation method:

[0282] The configuration interface offers options for pagination methods;

[0283] The system automatically captures the DOM selector when a user clicks the "Next Page" button on the page.

[0284] You can also manually enter the CSS selector or XPath for the pagination button;

[0285] Set the maximum number of page turns or the termination condition (such as detecting the "No more data" flag).

[0286] Configuration example:

[0287] {

[0288] "pagination":{

[0289] "type":"click",

[0290] "nextSelector":".next-page",

[0291] "maxPages":20

[0292] }

[0293] }

[0294] Step 4: Data Preview and Rule Testing

[0295] Objective: To preview the configuration effect in real time before the configuration is completed.

[0296] Implementation method:

[0297] According to the current rules, the first N rows of data are extracted and displayed in real time at the bottom of the configuration interface.

[0298] If pagination exists, you can manually test page turning and confirm the crawling effect;

[0299] It can display error messages, such as: Unable to locate pagination buttons, missing data, etc.

[0300] Technical points:

[0301] All preview results are not written to the database, but are rendered only in memory;

[0302] Use asynchronous data requests (Promise) to avoid page lag;

[0303] Supports exporting preview results as JSON / CSV sample files.

[0304] Step 5: Rule naming, saving, and reuse mechanism

[0305] Objective: To enable data collection rules to be named and saved for easy reuse.

[0306] Implementation method:

[0307] The configuration interface provides a "Rule Name" input box at the top;

[0308] After clicking the save button, the current rule structure will be saved to local storage (localStorage), synchronized with the browser, or uploaded to the server;

[0309] It supports three operation modes: "Create new rule", "Save overwrite", and "Import existing rule".

[0310] Code example:

[0311] {

[0312] "ruleId":"jd_table_202507",

[0313] "ruleName":"JD.com Product Data Collection",

[0314] "createdAt":"2025-07-22",

[0315] "rules":{...}

[0316] }

[0317] Scalability points:

[0318] Save as a "Frequently Used Data Collection Template" for the user account;

[0319] Set "rule tags" to facilitate categorized searching.

[0320] Step Six: Rule Format Validation and Task Submission Guidance

[0321] Objective: To verify the integrity of the rules and guide the user into the task execution process.

[0322] Implementation method:

[0323] When the "Submit Data Collection Task" button is clicked, the completeness of the rules is verified:

[0324] Are any valid tables selected?

[0325] Is at least one field configured?

[0326] Is the pagination setting reasonable?

[0327] After successful verification, a task confirmation window will be displayed to guide the user:

[0328] Naming tasks;

[0329] Choose the execution method (immediate / scheduled);

[0330] Submit the task to the scheduling engine.

[0331] Task information structure:

[0332] {

[0333] "taskId":"task_001",

[0334] "ruleRef":"jd_table_202507",

[0335] "executionType":"manual"

[0336] }

[0337] 3. Data Acquisition Module: The implementation steps of the automatic process execution and data acquisition module.

[0338] This module is responsible for automatically executing the entire process of webpage loading, element location, pagination, data extraction, data cleaning, and storage, based on the outputs of the first two modules (structure recognition and rule configuration). Its core capability lies in: accurately and efficiently acquiring structured webpage data through automated processes, adapting to the structural differences and changes of different websites.

[0339] Step 1: Task Scheduling and Rule Loading

[0340] Objective: To receive data collection tasks from the task system and load the corresponding rules.

[0341] Specific operations:

[0342] The scheduling system inputs task parameters, including the target URL, collection rule ID, and collection time;

[0343] The module parses the JSON file containing rules, reading field location rules, pagination methods, collection conditions, etc.

[0344] Initialize the log tracker to record the status of the crawling process (success / failure / exception, etc.).

[0345] Code example:

[0346] {

[0347] "taskId":"task_001",

[0348] "targetUrl":"https: / / example.com / table-list",

[0349] "rules":{

[0350] "fields":[...],

[0351] "pagination":{...},

[0352] "preprocessing":{...}

[0353] }

[0354] }

[0355] Step Two: Webpage Loading and Environment Preparation

[0356] Objective: To access the webpage and ensure that all page elements have finished loading.

[0357] Technical approach:

[0358] Use headless browsers (such as Puppeteer, Playwright, Selenium) to simulate real user browsing behavior;

[0359] Configure user agent and anti-scraping mechanism bypass strategies (such as random clicks, scrolling, and delays);

[0360] Wait for the target table element to finish loading (via CSS Selector, XPath, table keyword matching, etc.);

[0361] It supports setting timeout and retry mechanisms to avoid interrupting the process due to loading failure.

[0362] Code example:

[0363] await page.goto('https: / / example.com / table-list');

[0364] await page.waitForSelector('.data-table');

[0365] Step 3: Execution of Master Data Extraction Logic

[0366] Objective: To accurately extract structured data fields based on rules.

[0367] Implementation details:

[0368] Traverse the table rows in the webpage ( Limitations, support based on Extract each column according to the rules. The fields required in )

[0369] Each column field is selected by the configured selector to extract text, attribute values ​​(such as image src), or link href;

[0370] Supports field-level data processing: whitespace removal, unit conversion, regular expression extraction, data standardization, etc.

[0371] The data types of corresponding fields are automatically validated (date, number, etc.) to improve data quality.

[0372] Code example:

[0373]

[0374] Step 4: Pagination Processing and Merging of Multiple Pages

[0375] Objective: To capture multiple pages of data to ensure complete data collection.

[0376] Supported pagination types:

[0377] Click the "Next Page" button to turn the page;

[0378] Page turning is achieved by appending URL parameters (e.g., ?page=2);

[0379] Scroll-based pagination (supports pull-to-refresh sites);

[0380] Operating procedures:

[0381] After retrieving data from the current page, find the pagination control element according to the rules;

[0382] Perform the page turning action, wait for the page to refresh, and detect changes in the table;

[0383] The duplicate data extraction process continues until the maximum page count is reached or a termination condition is detected. Code example:

[0384] Step 5: Data Cleaning and Structure Preparation

[0385] Objective: To standardize and deduplicate the captured raw data.

[0386] Processing logic:

[0387] Null value handling: Clear blank rows and empty fields;

[0388] Type conversion: such as converting prices to floating-point numbers, standardizing date formats;

[0389] Regular expression extraction: such as extracting zip codes from addresses and brands from titles;

[0390] Duplicate data detection: Remove duplicates based on primary key fields (such as product ID, title, etc.).

[0391] Code example:

[0392]

[0393] Step Six: Data Storage and Export

[0394] Objective: To store the final data into a specified data interface or export format.

[0395] Supported storage methods:

[0396] Store in a database (such as MySQL, MongoDB, PostgreSQL);

[0397] Upload to the backend API interface (JSON format);

[0398] Export local files (CSV / Excel / JSON);

[0399] Send to a third-party platform (such as DingTalk, WeChat Work, Lark Robot, etc.) via middleware.

[0400] Data structure example (JSON):

[0401]

[0402]

[0403] 4. Implementation steps of the data storage and post-processing module.

[0404] This module is responsible for persistently storing the structured data extracted in the previous stage, and performing subsequent format cleaning, structure adjustment, visualization processing, and multi-channel export to provide users or systems with directly usable data results, supporting integration and reuse with low-code platforms.

[0405] Step 1: Structured Data Encapsulation

[0406] Objective: To encapsulate the collected data into standard data objects, unify the field format, and facilitate subsequent storage and processing.

[0407] Specific implementation:

[0408] Encapsulate it into a JSON or DataFrame structure, with each record having a uniform field key-value pair;

[0409] Add metadata fields, such as collection time, source URL, collection task ID, page index, etc.;

[0410] Type the data items (e.g., string, date, number, boolean, etc.);

[0411] Construct a mapping relationship between "standard model fields" and "original webpage fields" to improve reusability.

[0412] Code example:

[0413]

[0414] Step 2: Data Cleaning and Standardization

[0415] Objective: To improve data consistency and availability to meet the requirements of analysis or integration with other systems.

[0416] The processing content includes:

[0417] Deduplication: Remove duplicate records based on primary key fields (such as title, URL);

[0418] Standardized format:

[0419] Dates should be uniformly converted to ISO8601 format;

[0420] Amounts should be rounded to two decimal places.

[0421] Boolean fields are converted to 0 / 1;

[0422] Units should be standardized: for example, weight should be standardized to "grams" and area to "square meters";

[0423] Null value handling: Missing fields are uniformly filled with null or default values;

[0424] Regular expression transformation: such as extracting keywords, categories, and brand names from descriptions.

[0425] Step 3: Persistent Data Storage

[0426] Objective: To store the cleaned data on a stable and persistent data platform, supporting subsequent access, querying, and export.

[0427] Supported data storage types:

[0428] Implementation details:

[0429] Call the database driver or ORM tool (such as SQLAlchemy, Mongoose) to perform batch inserts;

[0430] Automatic table and database creation, with automatic field type matching;

[0431] Creating indexes (such as primary keys and date fields) can improve retrieval efficiency.

[0432] The data storage log records the operation results and exceptions.

[0433] Step 4: Exporting Results and Supporting Multiple Formats

[0434] Purpose: To provide standard data files for users to view, use, or import into other systems.

[0435] Export format support:

[0436] CSV / Excel (for office use);

[0437] JSON / XML (used for interfacing with system APIs);

[0438] Markdown / HTML tables (for report display);

[0439] Supports template export (such as uniform style Excel templates and general report templates);

[0440] Export logic:

[0441] Users select fields, time ranges, and export formats through the interface;

[0442] The module automatically converts the corresponding data into the target format;

[0443] For large data exports, use pagination or asynchronous compression and packaging.

[0444] Step 5: Building Reusable Interfaces for Low-Code Platforms

[0445] Objective: To provide a standardized interface for data referencing on low-code platforms.

[0446] Implementation method:

[0447] Provides a RESTful API interface that supports GET queries and POST submissions;

[0448] Supports GraphQL querying, enhancing front-end integration flexibility;

[0449] Configurable interface data structure, pagination parameters, and field filtering conditions;

[0450] Supports access control methods such as OAuth2 and Token;

[0451] Provides a "data source registration" function, allowing you to directly add sources in the low-code platform (e.g., collect tables → convert data sources → configure as data sources for table components);

[0452] Integrate data change notification mechanisms (such as Webhooks) to enable automatic refresh.

[0453] Step Six: Data Collection Result Traceability and Data Security Assurance

[0454] Objective: To ensure data quality, auditability, and user privacy security.

[0455] Implementation mechanism:

[0456] Data entries are bound to task IDs, user IDs, and source URLs for easy traceability.

[0457] Each storage action is automatically logged, and rollback is supported.

[0458] Sensitive information (such as account numbers and mobile phone numbers) is automatically encrypted or de-identified;

[0459] Configure backup strategies for the storage system to prevent data loss.

[0460] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A system for automatically acquiring page table data based on a workflow, characterized in that, include The structure recognition plugin automatically identifies the data table areas that can be collected on the webpage after the browser loads the webpage, and guides the user to configure data collection in a graphical form. The user-visual configuration interface is used to show users the results of webpage structure recognition and provide a WYSIWYG way to configure data collection rules, enabling automatic mapping and manual confirmation from webpage table structure to structured extraction parameters. The data capture module is responsible for automatically executing the entire process of webpage loading, element location, pagination navigation, data extraction, data cleaning, and storage, based on the output of the first two modules, namely the structure recognition plugin and the user visual configuration interface. The data post-processing module is responsible for persistently storing the structured data extracted in the previous stage, and performing subsequent format cleaning, structure adjustment, visualization processing, and multi-channel export to provide users or systems with directly usable data results, and supports integration and reuse with low-code platforms.

2. The system according to claim 1, characterized in that, The working steps of the structure recognition plugin are as follows: Step 1: Plugin Injection and Page Load Listening Develop browser extensions, define content_scripts configurations, and automatically inject scripts when a user visits a target website; Use window.onload, DOMContentLoaded, or MutationObserver to listen for when the webpage has finished loading; To prevent partial page refreshes from causing delayed data loading, a method is used to determine when the structure is ready to load by combining setInterval and DOM change listeners. Step 2: DOM Structure Parsing and Candidate Region Extraction Get all of the pages Simultaneously scanning structures with specific structural features 、 Tags; Extract features from its child nodes, such as text density, row and column structure, and whether there is a title, and construct a list of candidate data regions. Step 3: Data Structure Pattern Analysis and Field Identification Iterate through each row of the candidate region and extract the text content of each column; Compare whether the first row of content has the characteristics of a field title; Analyze the data types of all columns and construct a field type matrix; Step 4: Visual highlighting and numbering of candidate regions Use front-end technologies to add borders and semi-transparent masks to the candidate table area; Add a numbered label to the top left corner of each candidate area, which users can click to select; When the mouse hovers over the area, a preview of the first few rows of data extracted from that area is displayed. Step 5: User selection interaction and field confirmation guidance After the user clicks on the highlighted area, a configuration panel will pop up; The panel displays the table field names and data previews; Users can adjust the field mapping relationships; Supports operations such as renaming fields, enabling / disabling fields, and setting field types; Step Six: Generate Structure Rules and Initialize Data Acquisition Tasks Abstract the information from DOM selectors and field indexes into general rules; If pagination is enabled, a corresponding pagination logic description will be generated. Rules can be stored in the plugin's local storage, IndexedDB, or sent to the backend via API.

3. The system according to claim 2, characterized in that, Step two uses heuristic rules, including: The average number of child nodes per row is ≥2; The number of child nodes in all rows is basically the same; There is one line with a clear title word; Machine learning models can be introduced to predict the structure of historical structural samples; Step 3 extracts the result structure into a two-dimensional array format; Utilizes NLP for field name recognition, compatible with Chinese and multiple languages; Use regular expressions to determine the specific data types of date, amount, and phone number fields; In step five, the configuration interface can embed a drawer bar on the right side of the page to avoid redirection; All user configurations are saved in JSON rule format for easy data retrieval and use later.

4. The system according to claim 1, characterized in that, The user-visual configuration interface operates as follows: Step 1: Initializing and displaying the rule configuration panel After the user clicks on the identified table area or clicks the data collection configuration button, a configuration panel pops up on the side of the page; The panel can be built using HTML+CSS+JavaScript / Vue / React and adopts a responsive design; After the panel loads, it automatically fills in the data fields returned by the structure recognition module and displays a table preview. Step 2: Field Parameter Mapping and Field Editing Interaction Field name settings: Displays the title of each column, and supports renaming; Field Enable / Disable: Control whether to collect data for this field via a checkbox; Field type selection: Provides a drop-down list for users to select the field type; Field Preview: This feature displays sample data for the field from the first few rows, helping to determine its meaning. Step 3: Configure pagination rules The configuration interface offers options for pagination methods; The system automatically captures the DOM selector when a user clicks the "Next Page" button on the page. Alternatively, you can manually enter the CSS selector or XPath for the pagination button; Set the maximum number of page turns or the termination condition; Step 4: Data Preview and Rule Testing According to the current rules, the first N rows of data are extracted and displayed in real time at the bottom of the configuration interface. If pagination exists, you can manually test page turning and confirm the crawling effect; Error messages can be displayed; Step 5: Rule naming, saving, and reuse mechanism The configuration interface provides a rule name input box at the top; Clicking the save button will save the current rule structure to local storage, synchronize it with the browser, or upload it to the server. It supports three operation modes: creating new rules, saving overwrite rules, and importing existing rules. Step Six: Rule Format Validation and Task Submission Guidance Implementation method: When the "Submit Data Collection Task" button is clicked, the completeness of the rules is verified: Are any valid tables selected? Is at least one field configured? Is the pagination setting reasonable? After successful verification, a task confirmation window will be displayed to guide the user: Naming tasks; Select the execution method; Submit the task to the scheduling engine.

5. The system according to claim 4, characterized in that, In step one, front-end frameworks such as Ant Design, Element UI, and Vuetify can be used to quickly build the UI; the panel should support drag and drop and automatically adapt to the current screen size. In step two, tooltips provide field type descriptions and allow setting default values ​​or data cleaning rules. In step four, all preview results are not written to the database, but only rendered in memory; asynchronous data requests are used to avoid page lag; and preview results can be exported as JSON / CSV sample files. In step five, save it as a commonly used data collection template for the user account; Set rule tags to facilitate categorized searching.

6. The system according to claim 1, characterized in that, The data scraping module works as follows: Step 1: Task Scheduling and Rule Loading The scheduling system inputs task parameters, including the target URL, collection rule ID, and collection time; The module parses the JSON file containing rules, and reads the field location rules, pagination methods, and collection conditions. Initialize the log tracker to record the status of the crawling process; Step Two: Webpage Loading and Environment Preparation Use a headless browser to simulate real user browsing behavior; Configure user agent and anti-scraping mechanism bypass strategies; Wait for the target table element to finish loading; Supports setting timeout and retry mechanisms to avoid interrupting the process due to loading failure; Step 3: Execution of Master Data Extraction Logic Traverse the table rows in the webpage ( Label; Extract each column according to the rules. The fields required in the configuration; for each field, extract the text, attribute value, or link href according to the configured selector; Supports field-level data processing: whitespace removal, unit conversion, regular expression extraction, and data standardization; automatic type validation for corresponding fields improves data quality. Step 4: Pagination Processing and Merging of Multiple Pages After retrieving data from the current page, find the pagination control element according to the rules; Perform the page turning action, wait for the page to refresh, and detect changes in the table; The process of extracting duplicate data continues until the maximum number of pages is reached or a termination condition is detected. Step 5: Data Cleaning and Structure Preparation Null value handling: Clear blank rows and empty fields; Type conversion: Price converted to floating-point; Date format standardized; Regular expression extraction: Extracting zip codes from addresses and brands from titles; Duplicate data detection: Remove duplicates based on the primary key field; Step Six: Data Storage and Export Supported storage methods: Store in the database; Upload to the backend API interface; Exporting local files; Send to a third-party platform via middleware.

7. The system according to claim 1, characterized in that, The working steps of the data post-processing module are as follows: Step 1: Structured Data Encapsulation Encapsulate it into a JSON or DataFrame structure, with each record having a uniform field key-value pair; Add metadata fields; Type-mark the data items; Construct a mapping relationship between standard model fields and original webpage fields; Step 2: Data Cleaning and Standardization Deduplication: Remove duplicate records based on the primary key field; Standardized format; Step 3: Persistent Data Storage Call the database driver or ORM tool to perform batch inserts; Automatic table and database creation, with automatic field type matching; Creating indexes improves retrieval efficiency; Data storage logs record operation results and exceptions; Step 4: Exporting Results and Supporting Multiple Formats Export logic: Users select fields, time ranges, and export formats through the interface; The module automatically converts the corresponding data into the target format; For large data exports, use pagination or asynchronous compression and packaging. Step 5: Building Reusable Interfaces for Low-Code Platforms Provides a RESTful API interface that supports GET queries and POST submissions; Supports GraphQL querying, enhancing front-end integration flexibility; Configurable interface data structure, pagination parameters, and field filtering conditions; Supports OAuth2 and token-based access control; Provides a data source registration function, allowing you to add sources directly in the low-code platform; Integrate a data change notification mechanism to enable automatic refresh; Step Six: Data Collection Result Traceability and Data Security Assurance Implementation mechanism: Data entries are bound to task IDs, user IDs, and source URLs for easy traceability. Each storage action is automatically logged, and rollback is supported. Sensitive information is automatically encrypted or de-identified; Configure backup strategies for the storage system to prevent data loss.

8. The system according to claim 7, characterized in that, In step three, the supported data storage types are: relational databases, NoSQL databases, object storage, and local files. Step 4: Result Export and Multi-Format Support Export format support: CSV / Excel is used in office scenarios; JSON / XML is used to interface with system APIs; Markdown / HTML tables are used for report display; Template export is supported.

Citation Information

Patent Citations

  • Automatic webpage table data extraction method and device

    CN107992625A

  • Webpage data acquisition method and system based on Google browser plug-in

    CN110276041A

  • Webpage table data extraction method and device, computer equipment and storage medium

    CN113569170A

  • Method for realizing data capture in combination with browser plug-in

    CN120011620A

  • Server for Operating Website and Recording Medium

    KR1020070016528A