Visual webpage data crawling method and system based on large model
By combining Chromium-based automated browsing and screenshots with multimodal large-scale model analysis, the problem of traditional web crawlers having difficulty in capturing diverse web content is solved, highly accurate data extraction and logical reconstruction are achieved, and multiple storage formats and real-time data push are supported.
Patent Information
- Application Number
- CN202510744643.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional web crawlers have difficulty capturing diverse web content, especially key information such as graphical layout, dynamic rendering, and encrypted fonts. Existing solutions are unable to understand page structure and logic, resulting in incomplete and inaccurate data extraction.
It uses Chromium-based automated browsing and screenshots, combined with multimodal large model analysis and OCR recognition, and uses the vision-language aligned Transformer multimodal model to recognize table and graphic structures. It uses a large language model to reconstruct logic and format data, integrates distributed task scheduling and monitoring pipelines, and supports multiple storage and output formats.
It significantly improves the accuracy of data extraction from complex layouts and charts, avoids anti-crawling strategies, ensures the integrity and logic of output data, and supports multiple storage formats and real-time data push.
Smart Images

Figure CN120632181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network information extraction, and in particular to a method and system for crawling visual web page data based on a large model. Background Art
[0002] With the rapid development of Web 2.0 / 3.0 technologies, web page content presentation has become more diverse, including graphical layout, dynamic rendering, Canvas drawing or encrypted fonts. Traditional crawlers based on DOM parsing often find it difficult to capture key information and are easily intercepted by anti-crawling mechanisms.
[0003] Existing technologies include some solutions that use optical character recognition (OCR) to recognize imaged text, but these solutions are unable to understand the page structure and logic. Other solutions use deep learning for visual text extraction, but their recognition rates are low for complex fonts, multi-column layouts, or numbers embedded in charts, and they struggle to restore the semantic relationships of the original webpage. These deficiencies lead to incomplete and inaccurate data extraction, severely impacting subsequent analysis and utilization. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for crawling visual web page data based on a large model to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for crawling visual web page data based on a large model, comprising the following steps:
[0006] S1 Automated Browsing and Screenshots: Based on the Chromium browser, it performs automated browsing, uses a distributed crawler framework to manage concurrent tasks, simulates user interactions to trigger dynamic content loading, integrates verification codes and anti-crawling measures, supports both full-page and area screenshots, and manages screenshot files in a unified manner.
[0007] S2 image preprocessing: Slice the screenshots, perform rotation correction, perspective and stretch deformation processing, as well as image enhancement and denoising operations;
[0008] S3OCR recognition: Select and fine-tune the OCR engine, locate and separate text, and verify and correct recognition results;
[0009] S4 Multimodal Large Model Analysis: This model uses a Transformer multimodal model pre-trained with vision-language alignment. It inputs the original screenshot, OCR text bounding box coordinates, and a page DOM structure snapshot. The model then performs table and graphic structure recognition, logical association and hierarchical analysis, and defines the output format.
[0010] S5 result fusion: For the same area, it uses confidence threshold and weighted fusion algorithm, giving priority to high-precision text provided by OCR, resolves multi-source conflicts, and outputs unified results;
[0011] S6 Large Language Model Parsing: Design multiple sets of prompt templates, use the open domain large language model combined with instruction fine-tuning to perform logic reconstruction and formatting, and perform error detection and feedback on output structured data;
[0012] S7 data storage and output: supports multiple storage backends, provides multiple structured file format outputs or RESTful API interfaces, connects to message queues for real-time data push and streaming processing, and integrates monitoring pipelines and distributed task scheduling frameworks for system fault tolerance and monitoring.
[0013] Preferably, in step S1 of automated browsing and screenshots, the automated browsing architecture adopts a Chromium-based browser, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework adopts Scrapy+SeleniumPool; page behavior simulation includes simulating clicks, sliding, scrolling, and dragging user interaction operations, and setting a configurable waiting strategy; verification codes and anti-crawling countermeasures include integrating a human-computer collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and performing reverse analysis on common JS obfuscation and dynamic loading schemes; visual screenshot capture supports two modes: full-page screenshots and area screenshots, and unified naming and version management of screenshot files.
[0014] Preferably, in step S1 of automated browsing and screenshots, the automated browsing architecture adopts a Chromium-based browser, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework adopts Scrapy+SeleniumPool; page behavior simulation includes simulating clicks, sliding, scrolling, and dragging user interaction operations, and setting a configurable waiting strategy; verification codes and anti-crawling countermeasures include integrating a human-computer collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and performing reverse analysis on common JS obfuscation and dynamic loading schemes; visual screenshot capture supports two modes: full-page screenshots and area screenshots, and unified naming and version management of screenshot files.
[0015] Preferably, in step S4 multimodal large model analysis, the model architecture uses a Transformer multimodal model that has been pre-trained with vision-language alignment and is fine-tuned on a web page layout dataset; the input data includes the original screenshot, OCR text bounding box coordinates, and page DOM structure snapshots, and the page information is represented by splicing feature maps; the model outputs high-level semantic labels for each area, including table type, chart type, and ordinary text blocks, and can determine the position of row and column separators, infer cross-row / cross-column cells, and generate a two-dimensional table skeleton; each module is layered to generate a page structure tree, and dynamic components are annotated; the model output is a list of JSON objects, each object containing an area ID, semantic type, text or numerical content, and bounding box coordinates. For chart-type areas, additional metadata is output.
[0016] Preferably, in step S7 data storage and output, the storage solution supports multiple storage backends, including relational databases, document databases, distributed file systems, and provides transaction and batch write interfaces; the output format and interface support JSON, CSV, Parquet and multiple structured file formats, and can also directly provide a RESTful API interface for other systems to call, and connect to the message queue to achieve real-time data push and streaming processing; the system fault tolerance and monitoring integrate Prometheus+Grafana to monitor the performance indicators of each module of the pipeline, introduce a distributed task scheduling framework, and set a retry strategy and alarm mechanism.
[0017] A system for a large-scale model-based visual web page data crawling method includes the following modules:
[0018] Automated browsing and screenshot module: used for automated browsing in Chromium-based browsers, performing page behavior simulation, verification code and anti-crawling countermeasure processing, and visual screenshot capture;
[0019] Image preprocessing module: used to perform block slicing, rotation correction, perspective and stretching deformation processing on screenshots, as well as image enhancement and denoising operations;
[0020] OCR recognition module: used to select and fine-tune the OCR engine, locate and separate text, and verify and correct recognition results;
[0021] Multimodal Large Model Analysis Module: This module uses a Transformer multimodal model pre-trained with vision-language alignment, takes as input the original screenshot, OCR text bounding box coordinates, and a page DOM structure snapshot, performs table and graphic structure recognition, logical association and hierarchical analysis, and defines the output format.
[0022] Result fusion module: It is used to use confidence threshold and weighted fusion algorithm for the same area, give priority to the high-precision text provided by OCR, resolve multi-source conflicts, and output unified results;
[0023] Large language model parsing module: used to design multiple sets of prompt templates, using an open domain large language model combined with instruction fine-tuning to perform logic reconstruction and formatting, and to perform error detection and feedback on output structured data;
[0024] Data storage and output module: used to support multiple storage backends, provide multiple structured file format outputs or RESTful API interfaces, connect to message queues to implement real-time data push and streaming processing, and integrate monitoring pipelines and distributed task scheduling frameworks for system fault tolerance and monitoring.
[0025] Preferably, in the automated browsing and screenshot module, the automated browsing architecture adopts a Chromium-based browser, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework adopts Scrapy+SeleniumPool; page behavior simulation includes simulating clicks, sliding, scrolling, and dragging user interaction operations, and setting a configurable waiting strategy; verification codes and anti-crawling countermeasures include integrating a human-computer collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and performing reverse analysis on common JS obfuscation and dynamic loading schemes; visual screenshot capture supports both full-page screenshot and area screenshot modes, and unified naming and version management of screenshot files.
[0026] Preferably, in the multimodal large model analysis module, the model architecture uses a Transformer multimodal model that has been pre-trained for vision-language alignment and is fine-tuned on a web page layout dataset; the input data includes the original screenshot, OCR text bounding box coordinates, and page DOM structure snapshots, and the page information is represented by splicing feature maps; the model outputs high-level semantic labels for each area, including table type, chart type, and ordinary text blocks, and can determine the position of row and column separators, infer cross-row / cross-column cells, and generate a two-dimensional table skeleton; each module is layered to generate a page structure tree, and dynamic components are annotated; the model output is a list of JSON objects, each object contains an area ID, semantic type, text or numerical content, and bounding box coordinates. For chart-type areas, additional metadata is output.
[0027] Preferably, in the result fusion module, the fusion strategy is designed for the same area, and the high-precision text provided by OCR is given priority. If the OCR confidence is low or a special font is encountered, it is switched to the output of the multimodal model, and the confidence threshold and weighted fusion algorithm are defined to generate the final text and structure results; the multi-source conflict resolution module performs a secondary comparison and correction of the values or key fields that are inconsistent between the OCR and multimodal outputs through regular expressions and dictionaries, and outputs the conflict records and processing strategies to the log; the output unification module converts the fusion results into a predefined intermediate format according to a unified data mapping table to ensure consistency in downstream processing.
[0028] Preferably, in the data storage and output module, the storage solution supports multiple storage backends, including relational databases, document databases, and distributed file systems, and provides transaction and batch write interfaces; the output format and interface support multiple structured file formats such as JSON, CSV, and Parquet, and can also directly provide a RESTful API interface for other systems to call, and connect to the message queue to achieve real-time data push and streaming processing; the system fault tolerance and monitoring integrate Prometheus+Grafana to monitor the performance indicators of each module in the pipeline, introduce a distributed task scheduling framework, and set retry strategies and alarm mechanisms.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The large-model-based visual web page data crawling method and system proposed in this invention can handle various pictorial or dynamic web pages through screenshots and multimodal analysis without relying on the DOM structure; OCR and the multimodal large model complement each other, significantly improving the accuracy of data extraction from different fonts, complex layouts and charts; simulating real browsing behavior and combining manual verification can effectively circumvent common anti-crawling strategies; the large language model intelligently reorganizes page information according to prompts to ensure the integrity and logic of the output data. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0032] In order to clearly and completely describe the objectives and technical solutions of the present invention and make the advantages more clearly understood, the embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, not all of them, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0033] For example 1, please refer to Figure 1 The present invention provides a technical solution: a method for crawling visual web page data based on a large model, comprising the following steps:
[0034] 1. Step S1: Automated browsing and screenshots
[0035] 1.1 Automated Browsing Architecture
[0036] Choose a Chromium-based browser, such as Selenium + ChromeDriver or Puppeteer, to ensure compatibility with the latest web features. Use a distributed crawler framework (such as Scrapy + SeleniumPool) to manage concurrent tasks and ensure scalability for large-scale web crawling. For high-concurrency scenarios, introduce browser reuse and warm-up strategies to avoid wasted resources caused by frequent browser launches.
[0037] 1.2 Page behavior simulation
[0038] Perform simulated user interactions such as clicks, slides, scrolls, and drags on the page to trigger lazy loading or dynamic loading of content.
[0039] Set a configurable waiting strategy, combining explicit waits (WebDriverWait) with implicit waits to ensure that all target elements are rendered.
[0040] 1.3 Verification Code and Anti-Crawling Countermeasures
[0041] Integrate the human-machine collaborative verification interface to pass simple math and slider verification codes to manual or third-party verification code coding platforms.
[0042] Randomizes request headers, User-Agent, cookie pool, and proxy IP to reduce the probability of being blocked by websites. Performs reverse analysis of common JS obfuscation and dynamic loading schemes, pre-renders or executes key script code, and obtains the real DOM.
[0043] 1.4 Visual screenshot capture
[0044] Supports full-page screenshots and region screenshots, respectively for capturing long lists or key modules. In multi-browser parallel tasks, unified screenshot file naming and versioning are implemented for easy tracking.
[0045] 2. Step S2: Image preprocessing
[0046] 2.1 Block Slicing
[0047] Use image segmentation algorithms based on location and content (such as OpenCV's MSER or text region detection based on deep learning) to segment the screenshot. Identify different areas such as tables, lists, and icons and generate independent slices.
[0048] 2.2 Rotation Correction
[0049] For the detected text blocks, the main direction of the text is calculated (Hough line detection or minimum bounding rectangle direction) and the angle is corrected. For image blocks with tilt exceeding a threshold (such as 2°), an affine transformation is performed to ensure horizontal text input during OCR.
[0050] 2.3 Perspective and stretching deformation
[0051] For screenshot areas with perspective distortion, a four-point perspective transform (OpenCV getPerspectiveTransform) is used to correct it. A stretching or padding strategy is automatically selected based on the aspect ratio of the slice to fit the OCR window size.
[0052] 2.4 Image Enhancement and Denoising
[0053] Adaptive bilateral filtering and CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithms are used to enhance text contrast. Binarization and morphological opening and closing operations are performed on areas with complex backgrounds or low contrast to remove noise.
[0054] 3. Step S3: OCR recognition
[0055] 3.1OCR engine selection and fine-tuning
[0056] Prioritize using OCR engines that support multilingual and complex typography recognition (such as PaddleOCR, Tesseract LSTM++, and Microsoft Read API). Fine-tune the OCR model using a small number of annotated samples for the specific fonts of the target website to improve the accuracy of rare font recognition.
[0057] 3.2 Text Positioning and Line Breaks
[0058] Based on the bounding boxes returned by OCR, the text is separated and sorted from top to bottom and left to right. It handles continuous text, line breaks, and multi-column layouts, and merges paragraphs using a clustering algorithm based on connected component analysis.
[0059] 3.3 Result Verification and Error Correction
[0060] Recognition results are validated using a language model (e.g., based on a Trie tree or pre-trained language model) to automatically correct common typos. For specific formats like numbers, amounts, and dates, regular expressions or pattern matching are used for secondary validation to eliminate noise.
[0061] 4. Step S4: Multimodal Large Model Analysis
[0062] 4.1 Model Architecture and Input Format
[0063] We use a Transformer multimodal model pre-trained on vision-language alignment (e.g., CLIP, BLIP, or X-VLM) and fine-tune it on a web page layout dataset. The input data includes the original screenshot, the coordinates of the OCR text bounding box, and a JSON snapshot of the page's DOM structure. The page information is represented by concatenated feature maps.
[0064] 4.2 Table and Graphic Structure Recognition
[0065] The multimodal model outputs high-level semantic labels for each region, including table type (simple table, nested table), chart type (bar chart, line chart, etc.), and regular text blocks. For table structures, the model determines the location of row and column dividers, infers cells that span rows and columns, and generates a two-dimensional table skeleton.
[0066] 4.3 Logical Association and Hierarchical Analysis
[0067] Based on location and semantic tags, the model stratifies modules and generates a Page Object Tree. Dynamic components such as nested lists, menus, and pop-ups are annotated to ensure that interaction logic can be restored during downstream parsing.
[0068] 4.4 Output format definition
[0069] The model output is a list of JSON objects, each containing the region ID, semantic type, text or numeric content, and bounding box coordinates. For chart regions, additional metadata is output, such as legend labels, axis ranges, and a list of data points.
[0070] 5. Step S5: Result Fusion
[0071] 5.1 Integration Strategy Design
[0072] For the same region, the high-precision text provided by OCR is prioritized; if the OCR confidence is low or a unique font is encountered, the output of the multimodal model is switched. The confidence threshold and weighted fusion algorithm (linear weighting or Bayesian fusion) are defined to generate the final text and structure results.
[0073] 5.2 Multi-source conflict resolution
[0074] For values or key fields that are inconsistent between OCR and multimodal output, a secondary comparison and correction is performed using regular expressions and dictionaries. Conflict records and handling strategies are also output to the log for subsequent manual review and model retraining.
[0075] 5.3 Output Unification
[0076] Convert the fusion results into a predefined intermediate format (such as JSONSchema) according to a unified data mapping table to ensure consistency in downstream processing.
[0077] 6. Step S6: Large Language Model Parsing
[0078] 6.1Prompt Design and Template Parsing
[0079] Design multiple sets of prompt templates to output corresponding structured formats for different types of data, such as paragraph text, tables, lists, and charts. Use open-domain large language models (such as the GPT series or self-developed LLM) combined with instruction tuning to enhance parsing capabilities.
[0080] 6.2 Logical Reconstruction and Formatting
[0081] Based on the text block order and hierarchy information returned by Prompt, it reconstructs the logical order of paragraphs and restores the layout structure of titles, body text, comments, etc. It parses table content into a two-dimensional array, organizes chart data into key-value pairs, and supports exporting to CSV and Excel.
[0082] 6.3 Error Detection and Feedback
[0083] Perform schema validation on the structured data output by LLM, such as JSON validation, to ensure correct syntax and field types. Combined with a feedback loop, samples with parsing failures or inconsistent results are automatically retried or marked for manual intervention.
[0084] 7. Step S7: Data storage and output
[0085] 7.1 Storage Solution
[0086] Supports multiple storage backends: relational databases (MySQL / PostgreSQL), document databases (MongoDB), and distributed file systems (HDFS / S3). It provides transaction and batch write interfaces to ensure performance and data consistency for large-scale crawling tasks.
[0087] 7.2 Output Format and Interface
[0088] It supports multiple structured file formats such as JSON, CSV, and Parquet, and also provides a RESTful API for other systems to call. It integrates with message queues (Kafka / RabbitMQ) to enable real-time data push and streaming processing.
[0089] 7.3 System Fault Tolerance and Monitoring
[0090] Integrate Prometheus+Grafana to monitor the performance indicators of each module in the pipeline (latency, error rate, throughput); introduce a distributed task scheduling framework (such as Apache Airflow) and set up retry strategies and alarm mechanisms.
[0091] In the second embodiment, based on the first embodiment, a system for a large model-based visual webpage data crawling method is proposed, which includes the following modules:
[0092] Automated browsing and screenshot module: used for automated browsing in Chromium-based browsers, performing page behavior simulation, verification code and anti-crawling countermeasures processing, and visual screenshot capture; the automated browsing architecture uses Chromium-based browsers, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework uses Scrapy+SeleniumPool; page behavior simulation includes simulating clicks, sliding, scrolling, and dragging user interaction operations, and setting configurable waiting strategies; verification codes and anti-crawling countermeasures include integrating a human-computer collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and reverse analysis of common JS obfuscation and dynamic loading schemes; visual screenshot capture supports both full-page screenshot and area screenshot modes, and unified naming and version management of screenshot files.
[0093] Image preprocessing module: used to perform block slicing, rotation correction, perspective and stretching deformation processing on screenshots, as well as image enhancement and denoising operations;
[0094] OCR recognition module: used to select and fine-tune the OCR engine, locate and separate text, and verify and correct recognition results;
[0095] Multimodal large model analysis module: It is used to select the Transformer multimodal model pre-trained with vision-language alignment, input the original screenshot, OCR text bounding box coordinates, and page DOM structure snapshot, perform table and graphic structure recognition, logical association and hierarchical analysis, and define the output format; the model architecture uses the Transformer multimodal model pre-trained with vision-language alignment and is fine-tuned on a web page layout dataset; the input data includes the original screenshot, OCR text bounding box coordinates, and page DOM structure snapshot, and the page information is represented by splicing feature maps; the model outputs high-level semantic labels for each area, including table type, chart type, and ordinary text blocks, and can determine the position of row and column separators, infer cross-row / cross-column cells, and generate a two-dimensional table skeleton; each module is layered to generate a page structure tree and annotate dynamic components; the model output is a list of JSON objects, each object containing the area ID, semantic type, text or numerical content, and bounding box coordinates. For chart-type areas, additional metadata is output.
[0096] Result fusion module: used for the same area, using confidence threshold and weighted fusion algorithm, giving priority to the high-precision text provided by OCR, resolving multi-source conflicts, and outputting unified results; fusion strategy design: for the same area, giving priority to the high-precision text provided by OCR. If the OCR confidence is low or a special font is encountered, it switches to the output of the multimodal model, defines the confidence threshold and weighted fusion algorithm to generate the final text and structure results; the multi-source conflict resolution module performs secondary comparison and correction on inconsistent values or key fields between OCR and multimodal outputs through regular expressions and dictionaries, and outputs the conflict records and processing strategies to the log; the output unification module converts the fusion results into a predefined intermediate format according to a unified data mapping table to ensure consistent downstream processing.
[0097] Large language model parsing module: used to design multiple sets of prompt templates, using an open domain large language model combined with instruction fine-tuning to perform logic reconstruction and formatting, and to perform error detection and feedback on output structured data;
[0098] Data storage and output module: used to support multiple storage backends, provide multiple structured file format outputs or RESTful API interfaces, connect to message queues to achieve real-time data push and streaming processing, and integrate monitoring pipelines and distributed task scheduling frameworks for system fault tolerance and monitoring; storage solutions support multiple storage backends, including relational databases, document databases, and distributed file systems, and provide transaction and batch write interfaces; output formats and interfaces support multiple structured file formats such as JSON, CSV, and Parquet, and can also directly provide RESTful API interfaces for other systems to call, connecting to message queues to achieve real-time data push and streaming processing; system fault tolerance and monitoring integrate Prometheus+Grafana to monitor the performance indicators of each module in the pipeline, introduce a distributed task scheduling framework, and set retry strategies and alarm mechanisms.
[0099] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A visual webpage data crawling method based on a large model, characterized by: The following steps are involved: S1 Automated Browsing and Screenshots: Based on the Chromium browser, it performs automated browsing, uses a distributed crawler framework to manage concurrent tasks, simulates user interactions to trigger dynamic content loading, integrates verification codes and anti-crawling measures, supports both full-page and area screenshots, and manages screenshot files in a unified manner. S2 image preprocessing: Slice the screenshots, perform rotation correction, perspective and stretch deformation processing, as well as image enhancement and denoising operations; S3OCR recognition: Select and fine-tune the OCR engine, locate and separate text, and verify and correct recognition results; S4 Multimodal Large Model Analysis: This model uses a Transformer multimodal model pre-trained with vision-language alignment. It inputs the original screenshot, OCR text bounding box coordinates, and a page DOM structure snapshot. The model then performs table and graphic structure recognition, logical association and hierarchical analysis, and defines the output format. S5 result fusion: For the same area, it uses confidence threshold and weighted fusion algorithm, giving priority to high-precision text provided by OCR, resolves multi-source conflicts, and outputs unified results; S6 Large Language Model Parsing: Design multiple sets of prompt templates, use the open domain large language model combined with instruction fine-tuning to perform logic reconstruction and formatting, and perform error detection and feedback on output structured data; S7 data storage and output: supports multiple storage backends, provides multiple structured file format outputs or RESTful API interfaces, connects to message queues for real-time data push and streaming processing, and integrates monitoring pipelines and distributed task scheduling frameworks for system fault tolerance and monitoring.
2. The method for crawling visual webpage data based on a large model according to claim 1, characterized in that: In step S1, automated browsing and screenshots are performed. The automated browsing architecture uses a Chromium-based browser, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework uses Scrapy+SeleniumPool. Page behavior simulation includes simulating clicks, slides, scrolls, and dragging user interactions, and setting configurable waiting strategies. Verification codes and anti-crawl countermeasures include integrating a human-machine collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and reverse analysis of common JS obfuscation and dynamic loading schemes; visual screenshot capture supports both full-page screenshot and area screenshot modes, and unified naming and version management of screenshot files.
3. The method for crawling webpage data based on a large model according to claim 2, characterized in that: In step S2 image preprocessing, block slicing uses an image segmentation algorithm based on position and content to divide the screenshot into blocks, identify different areas separately and generate independent slices; rotation correction calculates the main direction of the detected text blocks and performs angle correction; perspective and stretching deformation performs four-point perspective transformation correction on the screenshot area with perspective distortion, and automatically selects stretching or white-filling strategy according to the aspect ratio of the slice; image enhancement and denoising apply adaptive bilateral filtering and CLAHE algorithm to improve text contrast, and perform binarization and morphological opening and closing operations on areas with complex backgrounds or low contrast.
4. The method for crawling webpage data based on a large model according to claim 1, characterized in that: In step S4, the multimodal large model analysis, the model architecture uses a Transformer multimodal model pre-trained with vision-language alignment and fine-tuned on a web page layout dataset. The input data includes the original screenshot, the coordinates of the OCR text bounding box, and a snapshot of the page DOM structure. The page information is represented by splicing feature maps. The model outputs high-level semantic labels for each region, including table type, chart type, and ordinary text blocks. It can also determine the position of row and column dividers, infer cross-row / cross-column cells, and generate a two-dimensional table skeleton. Each module is layered to generate a page structure tree and annotate dynamic components. The model output is a list of JSON objects, each of which contains the region ID, semantic type, text or numerical content, and bounding box coordinates. For chart-type regions, additional metadata is output.
5. The method for crawling visual webpage data based on a large model according to claim 1, characterized in that: In step S7 data storage and output, the storage solution supports multiple storage backends, including relational databases, document databases, and distributed file systems, and provides transaction and batch write interfaces; the output format and interface support multiple structured file formats such as JSON, CSV, and Parquet, and can also directly provide a RESTful API interface for other systems to call, connecting to the message queue to achieve real-time data push and streaming processing; system fault tolerance and monitoring integrate Prometheus+Grafana to monitor the performance indicators of each module in the pipeline, introduce a distributed task scheduling framework, and set retry strategies and alarm mechanisms.
6. A system for the large model-based visual webpage data crawling method according to claim 5, characterized in that: Includes the following modules: Automated browsing and screenshot module: used for automated browsing in Chromium-based browsers, performing page behavior simulation, verification code and anti-crawling countermeasure processing, and visual screenshot capture; Image preprocessing module: used to perform block slicing, rotation correction, perspective and stretching deformation processing on screenshots, as well as image enhancement and denoising operations; OCR recognition module: used to select and fine-tune the OCR engine, locate and separate text, and verify and correct recognition results; Multimodal Large Model Analysis Module: This module uses a Transformer multimodal model pre-trained with vision-language alignment, takes as input the original screenshot, OCR text bounding box coordinates, and a page DOM structure snapshot, performs table and graphic structure recognition, logical association and hierarchical analysis, and defines the output format. Result fusion module: It is used to use confidence threshold and weighted fusion algorithm for the same area, give priority to the high-precision text provided by OCR, resolve multi-source conflicts, and output unified results; Large language model parsing module: used to design multiple sets of prompt templates, using an open domain large language model combined with instruction fine-tuning to perform logic reconstruction and formatting, and to perform error detection and feedback on output structured data; Data storage and output module: used to support multiple storage backends, provide multiple structured file format outputs or RESTful API interfaces, connect to message queues to implement real-time data push and streaming processing, and integrate monitoring pipelines and distributed task scheduling frameworks for system fault tolerance and monitoring.
7. A system according to claim 6, characterized in that: In the automated browsing and screenshot module, the automated browsing architecture uses a Chromium-based browser, such as Selenium+ChromeDriver or Puppeteer, and the distributed crawler framework uses Scrapy+SeleniumPool. Page behavior simulation includes simulating clicks, slides, scrolling, and dragging user interactions, and setting configurable waiting strategies. Verification codes and anti-crawl countermeasures include integrating a human-machine collaborative verification interface, randomizing request headers, User-Agent, Cookie pool, and proxy IP, and reverse analysis of common JS obfuscation and dynamic loading schemes; visual screenshot capture supports both full-page screenshot and area screenshot modes, and unified naming and version management of screenshot files.
8. A system according to claim 6, characterized in that: In the multimodal large model analysis module, the model architecture uses a Transformer multimodal model pre-trained with vision-language alignment and fine-tuned on a web page layout dataset. The input data includes the original screenshot, the coordinates of the OCR text bounding box, and a snapshot of the page DOM structure. The page information is represented by concatenated feature maps. The model outputs high-level semantic labels for each region, including table type, chart type, and regular text blocks. It can also determine the location of row and column dividers, infer cross-row / cross-column cells, and generate a two-dimensional table skeleton. Each module is layered to generate a page structure tree and annotate dynamic components. The model output is a list of JSON objects, each of which contains the region ID, semantic type, text or numerical content, and bounding box coordinates. For chart-type regions, additional metadata is output.
9. A system according to claim 6, characterized in that: In the result fusion module, the fusion strategy is designed for the same area, giving priority to the high-precision text provided by OCR. If the OCR confidence is low or a special font is encountered, it switches to the output of the multimodal model, defines the confidence threshold and weighted fusion algorithm to generate the final text and structure results; the multi-source conflict resolution module performs a secondary comparison and correction of inconsistent values or key fields between OCR and multimodal outputs through regular expressions and dictionaries, and outputs the conflict records and processing strategies to the log; the output unification module converts the fusion results into a predefined intermediate format according to a unified data mapping table to ensure consistency in downstream processing.
10. A system according to claim 6, characterized in that: In the data storage and output module, the storage solution supports multiple storage backends, including relational databases, document databases, and distributed file systems, and provides transaction and batch write interfaces; the output format and interface support multiple structured file formats such as JSON, CSV, and Parquet, and can also directly provide a RESTful API interface for other systems to call, and connect to the message queue to achieve real-time data push and streaming processing; system fault tolerance and monitoring integrate Prometheus+Grafana to monitor the performance indicators of each module in the pipeline, introduce a distributed task scheduling framework, and set retry strategies and alarm mechanisms.
Citation Information
Patent Citations
Dynamic chart page data crawling method and apparatus, terminal and storage medium
CN108595583A
Multi-modal Action Transform model and intelligent task execution method thereof
CN119494078A
Power grid field table recognition and information extraction method based on large model
CN119942574A
Document information extraction method, device and system and storage medium
CN119942576A
Cited By
Policy text multi-modal acquisition and evolution graph analysis method and policy text multi-modal acquisition and evolution graph analysis system
CN122174844A