Method and system for automatically extracting list data, storage medium and electronic equipment
By analyzing the web page structure, filtering and determining the target data list, and extracting the data to output in a preset format, the problem that existing tools cannot accurately identify complex web page target lists and data formats is solved, and efficient and accurate data extraction and processing are achieved.
Patent Information
- Application Number
- CN202411946475.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing automatic data extraction tools cannot accurately identify target lists in complex web pages, the data format is not standardized, and the system is poor, and it is susceptible to changes in web page structure.
By obtaining web page content, analyzing the web page structure to identify the geometric information and path information of clickable elements, forming a preliminary list structure, filtering vertically arranged lists, comprehensively analyzing geometric information and user filtering conditions to determine the target data list, and extracting the data to output in a preset format.
It realizes the rapid and accurate extraction of web page target list data, improves data extraction efficiency, ensures data accuracy and standardization, enhances data processing capabilities, and improves user experience.
Smart Images

Figure CN120067477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a method, system, storage medium and electronic device for automatically extracting list data. Background Art
[0002] With the rapid development of Internet technology, web pages, as one of the main carriers of information dissemination, contain a vast amount of data resources. These data resources are often presented in the form of lists, such as product information, news lists, user comments, etc. However, traditional data extraction methods mostly rely on manual operations, which are not only inefficient but also error-prone. Therefore, it is particularly important to develop a system and method that can automatically extract web page list data.
[0003] Currently, there are already some automatic data extraction tools on the market, but most of them have the following problems: First, they cannot accurately identify the target lists in web pages, especially when the web page structure is complex and changeable; second, the extracted data format is not standardized and is difficult to be directly used for subsequent data processing and analysis; third, the system stability is poor and is easily affected by changes in the web page structure. To solve the above problems, we propose a method, system, storage medium and electronic device for automatically extracting list data. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a method, system, storage medium and electronic device for automatically extracting list data, which solves the problems raised in the background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method, system, storage medium and electronic device for automatically extracting list data, specifically including the following steps:
[0006] S1. Obtain the web page content, parse the web page content to identify and extract the geometric information and path information of all clickable elements in the web page, and at the same time obtain the overall geometric information of the web page;
[0007] S2. According to the hierarchical relationship of the path information, classify and aggregate all clickable elements to form a preliminary list structure;
[0008] S3. Analyze the geometric information of the elements in each preliminary list structure, and screen out the vertically arranged lists as candidate target data lists;
[0009] S4. Comprehensively analyze the geometric information of the candidate target data list, the geometric information of each element in the list, and the overall geometric information of the web page to determine the target data list;
[0010] S5. Extract the specific information of each element in the target data list and output the extracted information in a preset format.
[0011] Preferably, in step S1, the geometric information includes the position and size of the element, and the path information is the position of the element in the document object model tree.
[0012] Preferably, in step S2, the classification and aggregation are based on the hierarchical similarity of the path information, and the elements with the same or similar hierarchical paths are grouped into the same class.
[0013] Preferably, in step S3, the vertical list is filtered based on the vertical position relationship of the elements in the list. When the vertical distance between adjacent elements in the list is less than the preset threshold, the list is considered to be vertically arranged.
[0014] Preferably, in step S4, determining the target data list also considers the filtering conditions preset by the user, such as the list title and specific element identifiers in the list.
[0015] Preferably, the preset format in step S5 includes but is not limited to JSON, CSV, or XML, and the output information may include but is not limited to text content, link addresses, and picture URLs.
[0016] The present invention also discloses a system for automatically extracting list data, including a data acquisition module, a data processing module, and a data output module. The output end of the data acquisition module is electrically connected to the input end of the data processing module, and the output end of the data processing module is electrically connected to the data output module. The data acquisition module is used to acquire web page content and parse it to extract element information and page information. The data acquisition module includes a web crawler unit and an element parsing unit. The data processing module includes a classification and aggregation unit, a list filtering unit, a target determination unit, and a data extraction unit. The data output module includes a format conversion unit and an export interface unit. The data processing module is used to output the extracted data in a preset format.
[0017] Preferably, the web crawler unit is responsible for accessing the target web page and acquiring the HTML source code. The element parsing unit is used to parse the HTML source code to extract element information and page information. The classification and aggregation unit performs classification and aggregation according to the element path hierarchy. The list filtering unit is used to analyze the geometric information of the elements to filter the vertical list. The target determination unit is used to comprehensively judge and determine the target data list. The data extraction unit is used to extract the specific information of each element in the target data list. The format conversion unit converts the extracted data into a standard format. The export interface unit is used to provide a data export interface and support multiple data export methods.
[0018] The present invention also discloses a storage medium for automatically extracting list data, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 6 is implemented.
[0019] The present invention also discloses an electronic device for automatically extracting list data, including a processor and a memory. Computer program instructions are stored in the memory. When the processor executes the computer program instructions, the method described in any one of claims 1 to 6 is implemented.
[0020] Beneficial effects
[0021] The present invention provides a method, a system, a storage medium and an electronic device for automatically extracting list data. Compared with the prior art, the following beneficial effects are achieved:
[0022] (1) Through automated means, target list data can be quickly and accurately extracted from web pages, avoiding the time-consuming, laborious and error-prone problems in traditional manual extraction methods. The automated processing method significantly improves the efficiency of data extraction, enabling enterprises or individuals to be more efficient and accurate when processing a large amount of web page data.
[0023] (2) The present invention not only realizes the automatic extraction of data, but also classifies, filters and organizes the extracted data through complex algorithms and logics, further enhancing the data processing ability. This helps users better understand and utilize the data, providing strong support for subsequent data analysis, decision-making and other tasks.
[0024] (3) For users who need to frequently extract data from web pages, the application of the present invention can undoubtedly greatly improve their user experience. Users no longer need to manually copy and paste or write complex script programs. They can obtain the required data with simple operations, greatly saving time and effort. At the same time, this patent also provides flexible data export methods, facilitating users to import the data into other data processing software for further analysis and processing. Brief description of the drawings
[0025] Figure 1 It is a system block diagram of the present invention;
[0026] Figure 2 It is a schematic diagram of the data acquisition module of the present invention;
[0027] Figure 3 It is a schematic diagram of the data processing module of the present invention;
[0028] Figure 4 It is a schematic diagram of the data output module of the present invention.
[0029] In the figure: 01, data acquisition module; 02, data processing module; 03, data output module; 011, web crawler unit; 012, element parsing unit; 021, classification and aggregation unit; 022, list screening unit; 023, target determination unit; 024, data extraction unit; 031, format conversion unit; 032, export interface unit. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] The embodiments of the present invention provide three technical solutions, which specifically include the following embodiments:
[0032] Embodiment 1
[0033] A method, system, storage medium and electronic device for automatically extracting list data specifically include the following steps:
[0034] S1. Install a Python development environment on a computer and configure relevant libraries (such as requests, BeautifulSoup, etc.) to support web crawling and HTML parsing functions. Design and implement the relevant code of the data acquisition module, data processing module and data output module. Obtain the web page content, parse the web page content to identify and extract the geometric information and path information of all clickable elements in the web page, and at the same time obtain the overall geometric information of the web page. Use web crawling technology to access the target web page, obtain the HTML source code of the web page, parse the HTML source code, extract the geometric information (such as position, size) and path (such as position in the DOM tree) of all clickable elements in the web page, and at the same time, obtain the overall geometric information of the web page, such as page width, height, etc.;
[0035] S2. According to the path hierarchy relationship of the elements, use intelligent algorithms to classify them in detail. In this process, the system can accurately identify elements with the same or similar hierarchical paths and classify them into the same category. Subsequently, perform element aggregation within the same category, and initially construct a list structure based on their mutual relevance and layout characteristics, laying a solid foundation for subsequent data screening and extraction.
[0036] S3. To screen out the target data list from numerous preliminary lists, the system deeply analyzes the geometric information of each list element, especially their vertical positional relationship. By setting a reasonable threshold, when the vertical distance between adjacent elements in the list is less than this threshold, it is determined that the list is vertically arranged and retained as a candidate target data list. This step effectively eliminates the interference of non-target data and improves the accuracy of data extraction.
[0037] S4. After obtaining the candidate target data list, the system further comprehensively analyzes the geometric information of the vertical list, the geometric information of each element in the list, and the overall geometric information of the web page, while considering the filtering conditions preset by the user (such as specific list titles, element identifiers, etc.) to comprehensively determine the true target data list. Once the target data list is locked, the system immediately starts the information extraction program to accurately capture the specific information of each element in the list (such as text content, link address, picture URL, etc.) to ensure the integrity and accuracy of the data.
[0038] S5. The system formats the extracted data in a specification format specified by the user (such as JSON, CSV, XML, etc.) and outputs it to the specified location. At the same time, a convenient data export interface is provided to support the user to easily import the extracted data into various data processing and analysis software, greatly improving the efficiency and value of data utilization.
[0039] In step S1, the geometric information includes the position and size of the element, and the path information is the position of the element in the document object model tree.
[0040] In step S2, the classification and aggregation are based on the hierarchical similarity of the path information, and the elements with the same or similar hierarchical paths are grouped into the same category.
[0041] In step S3, screening the vertically arranged list is based on the vertical positional relationship of the elements in the list. When the vertical distance between adjacent elements in the list is less than the preset threshold, the list is considered to be vertically arranged.
[0042] In step S4, determining the target data list also considers the filtering conditions preset by the user, such as the list title and specific element identifiers in the list.
[0043] In step S5, the preset formats include but are not limited to JSON, CSV, or XML, and the output information may include but is not limited to text content, link address, picture URL.
[0044] Embodiment 2
[0045] Based on Embodiment 1, refer to Figures 1 - 4As shown in the figure, the present invention also discloses a system for automatically extracting list data, including a data acquisition module 01, a data processing module 02, and a data output module 03. The output end of the data acquisition module 01 is electrically connected to the input end of the data processing module 02, and the output end of the data processing module 02 is electrically connected to the data output module 03. The data acquisition module 01 is used to acquire web page content and parse it to extract element information and page information. The data acquisition module 01 includes a web crawler unit 011 and an element parsing unit 012. The data processing module 02 includes a classification and aggregation unit 021, a list screening unit 022, a target determination unit 023, and a data extraction unit 024. The data output module 03 includes a format conversion unit 031 and an export interface unit 032. The data processing module 02 is used to output the extracted data in a preset format.
[0046] The web crawler unit 011 is responsible for accessing the target web page and obtaining the HTML source code. The element parsing unit 012 is used to parse the HTML source code to extract element information and page information. The classification and aggregation unit 021 performs classification and aggregation according to the element path hierarchy. The list screening unit 022 is used to analyze the element geometric information and screen the vertical list. The target determination unit 023 is used to comprehensively judge and determine the target data list. The data extraction unit 024 is used to extract the specific information of each element in the target data list. The format conversion unit 031 converts the extracted data into a standard format. The export interface unit 032 is used to provide a data export interface and support multiple data export methods.
[0047] Embodiment 3
[0048] On the basis of Embodiment 2, the present invention also discloses a storage medium for automatically extracting list data, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method according to any one of Claims 1 to 6 is implemented. The storage medium can be any type of computer-readable medium, such as a hard disk, a USB flash drive, an SD card, etc.
[0049] The present invention also discloses an electronic device for automatically extracting list data, including a processor and a memory. Computer program instructions are stored in the memory to implement the function of automatically extracting list data. The electronic device can be a device with data processing capabilities such as a computer, a server, a smart phone, etc. When the processor executes the computer program instructions, the method according to any one of Claims 1 to 6 is implemented.
[0050] Meanwhile, the content not described in detail in this specification belongs to the prior art well known to those skilled in the art.
[0051] The above has described the embodiments of the invention in detail, but the above content is only the preferred embodiments of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made in accordance with the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. A method for automatically extracting list data, characterized in that: The specific steps include: S1. Obtain web page content, parse the web page content to identify and extract geometric information and path information of all clickable elements in the web page, and simultaneously obtain overall geometric information of the web page; S2. Classify and aggregate all clickable elements according to the hierarchical relationship of the path information to form a preliminary list structure; S3, analyzing the geometric information of the elements in each preliminary list structure, and selecting the vertically arranged lists as candidate target data lists; S4, comprehensively analyzing the geometric information of the candidate target data list, the geometric information of each element in the list, and the overall geometric information of the web page to determine the target data list; S5. Extract specific information of each element in the target data list, and output the extracted information in a preset format.
2. The method for automatically extracting list data according to claim 1, characterized in that: The geometric information in step S1 includes the position and size of the element, and the path information is the position of the element in the document object model tree.
3. The method for automatically extracting list data according to claim 1, characterized in that: The classification aggregation in step S2 is performed based on the hierarchical similarity of the path information, and the elements with the same or similar hierarchical paths are classified into the same category.
4. The method for automatically extracting list data according to claim 1, characterized in that: The screening of the vertically arranged list in step S3 is performed based on the vertical position relationship of the elements in the list. When the vertical distance between adjacent elements in the list is less than a preset threshold, the list is considered to be vertically arranged.
5. The method for automatically extracting list data according to claim 1, characterized in that: The determination of the target data list in step S4 also takes into account the screening conditions preset by the user, such as the list title and the identifier of a specific element in the list.
6. The method for automatically extracting list data according to claim 1, characterized in that: The preset format in step S5 includes but is not limited to JSON, CSV or XML, and the output information may include but is not limited to text content, link address, and image URL.
7. A system for automatically extracting list data, characterized in that: The invention comprises a data acquisition module (01), a data processing module (02), and a data output module (03); the output end of the data acquisition module (01) is electrically connected to the input end of the data processing module (02); the output end of the data processing module (02) is electrically connected to the data output module (03); the data acquisition module (01) is used to acquire web page content and parse it to extract element information and page information; the data acquisition module (01) comprises a web crawler unit (011) and an element parsing unit (012); the data processing module (02) comprises a classification aggregation unit (021), a list screening unit (022), a target determination unit (023), and a data extraction unit (024); the data output module (03) comprises a format conversion unit (031) and an export interface unit (032); the data processing module (02) is used to output the extracted data in a preset format.
8. The system for automatically extracting list data according to claim 7, characterized in that: The web crawler unit (011) is responsible for accessing the target web page and obtaining the HTML source code; the element parsing unit (012) is used to parse the HTML source code and extract element information and page information; the classification aggregation unit (021) performs classification aggregation according to the element path level; the list screening unit (022) is used to analyze the element geometry information and screen the vertical list; the target determination unit (023) is used to determine the target data list through comprehensive judgment; the data extraction unit (024) is used to extract the specific information of each element in the target data list; the format conversion unit (031) converts the extracted data into a standard format; and the export interface unit (032) is used to provide a data export interface and support multiple data export methods.
9. A storage medium for automatically extracting list data, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. An electronic device for automatically extracting list data, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, the method according to any one of claims 1 to 6 is implemented.