Method and System for Extracting Main Text of Ethnic Group Web Pages Based on Statistical Rules

The statistical rule-based method addresses the challenge of extracting web page text from diverse HTML structures by using a redundancy dictionary to filter and preserve text sequence, improving extraction precision and efficiency.

CN115510307BActive Publication Date: 2025-07-15SHANDONG EVAYINFO TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211200790.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-07-15
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract web page text with irregular HTML structure, resulting in increased computing costs and affected parsing effects.

Method used

The homepage text extraction method based on statistical rules is adopted. By establishing a deduplication dictionary and setting thresholds, repeating strings are removed, text sequence structure is preserved, starting and ending positions are positioned, and text text is output.

Benefits of technology

It improves the accuracy and efficiency of web page text extraction, is suitable for different forms of web pages, and the extracted content is highly readable and has a coherent word order, and is suitable for deep learning text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510307B_ABST
    Figure CN115510307B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for extracting the main text of ethnic group web pages based on statistical rules, which obtains a set of web pages to be processed in the form of a web page ethnic group and gets a web page ethnic group list; traverses the web page ethnic group list, extracts the original HTML code of each web page, and forms an HTML code list; traverses the HTML code list, extracts all the text content in each web page, and according to the HTML structure, converts the long texts of all web pages into a list of short text strings and retains the text order; wherein, each list of short text strings belongs to the text list set of the entire web page ethnic group; traverses the text list set and locates the starting position and the ending position for each list of short text strings; selects the text from the starting position to the ending position and outputs the main text list; the present invention does not require manual participation and special rules, can extract web page texts in different forms, and greatly improves the extraction accuracy and extraction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval, and particularly to a method and system for extracting ethnic group web page text based on statistical rules. Background Art

[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] The extraction of web page text content is the preliminary work of web page content parsing. In order to better display web pages, they usually contain a lot of information irrelevant to the text, such as website navigation lists, website titles, copyright marks, and so on. The above-mentioned information can provide a better browsing experience, but this information is useless and interfering to the web page parsing system. If the web page information saved in HTML form cannot be preprocessed (i.e., extracting the text of the web page) in the early stage, the text parsing system will face a large amount of useless and messy text, and the length of this text can even exceed the text of the main body itself, which not only increases the calculation cost but also has a certain impact on the parsing effect.

[0004] Regarding the work of web page text extraction, there have been rich developments at the present stage: Patent No. CN111966901A discloses a method, system, device and storage medium for extracting policy-related web page text, which determines the position of the text through HTML source code; Patent No. CN110795933A discloses a method and device for identifying and processing web page text, which uses the number of words in a text block to determine the position of the text; Patent No. CN109948089A discloses a method and device for extracting web page text, which uses a neural network method to extract web page text; Patent No. CN105183801A provides a method and device for extracting web page text, which uses special structures and artificial rules in HTML to determine the position of the web page text.

[0005] The inventors found that the above methods, whether based on rules or on mathematical statistics, have an inherent disadvantage that they are only applicable to specific web pages or web pages with standardized HTML structures; although the above methods are very effective for extracting the text of web pages of some portal websites, in actual situations, the HTML structures of most web pages on the Internet are complex and non-standard, and effective extraction often cannot be achieved. Summary of the Invention

[0006] In order to solve the deficiencies of the prior art, the present invention provides a method and system for extracting ethnic group web page text based on statistical rules, which do not require manual participation and special rules, can extract web page texts in different forms, and can retain the basic order structure of the web page text, greatly improving the extraction accuracy and extraction efficiency.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In the first aspect of the present invention, a method for extracting ethnic group web page text based on statistical rules is provided.

[0009] A method for extracting ethnic group web page text based on statistical rules includes the following processes:

[0010] Obtain a set of web pages to be processed in the form of a web page ethnic group to obtain a web page ethnic group list;

[0011] Traverse the web page ethnic group list, extract the original HTML code of each web page, and form an HTML code list;

[0012] Traverse the HTML code list, extract all text contents in each web page, and according to the HTML structure, convert the long texts of all web pages into a list of short text strings, and retain the text order; among them, each list of short text strings belongs to the text list set of the entire web page ethnic group;

[0013] Establish a deduplication dictionary, where the keys of the deduplication dictionary are text strings, and the values of the deduplication dictionary are the number of times the string appears in the entire web page ethnic group. Traverse the text list set and remove duplicate strings from each list of short text strings to obtain a deduplicated list of short text strings, and fill the deduplication dictionary accordingly;

[0014] If the value corresponding to the string in the deduplication dictionary is greater than the set threshold, the string is excluded, otherwise the string is retained;

[0015] Traverse the text list set and locate the start position and end position for each list of short text strings;

[0016] Select the text from the start position to the end position and output the text list of the main body.

[0017] As an optional implementation manner, the web page ethnic group is delimited based on the content of the website navigation bar.

[0018] As an optional implementation manner, the traversed and filled deduplication dictionary represents the number of times each string appears in the entire web page ethnic group.

[0019] As an optional implementation manner, traversing the text list set and locating the start position for each list of short text strings includes:

[0020] Traverse the list of short text strings from the beginning, and for each string in the list of short text strings Look up the number of occurrences in the deduplication dictionary until found Then the j position is the start position of the main body, where t is the set threshold.

[0021] As an alternative implementation, traverse the text list set and locate the end position for each short text string list, including:

[0022] Traverse the short text string list from the end, for each string in the short text string list Look up the occurrence count in the deduplication dictionary until found Then the position at j is the end position of the main text, where t is a set threshold.

[0023] As an alternative implementation, traverse the text list set and locate the start position and end position for each short text string list, including:

[0024] Traverse the short text string list, remove the strings whose corresponding values in the deduplication dictionary are greater than the set threshold, and retain the original string order;

[0025] Select the position of the first string in the original text of the deduplicated string data group as the start position of the main text, and select the position of the last string in the original text of the deduplicated string data group as the end position of the main text.

[0026] As an alternative implementation, when the main text needs to be output in text form, use special delimiters for text segmentation.

[0027] The second aspect of the present invention provides a system for extracting the main text of ethnic group web pages based on statistical rules.

[0028] A system for extracting the main text of ethnic group web pages based on statistical rules, including:

[0029] A web page ethnic group list generation module, configured to: obtain a group of web pages to be processed in the form of a web page ethnic group, and obtain a web page ethnic group list;

[0030] An HTML code list generation module, configured to: traverse the web page ethnic group list, extract the original HTML code of each web page, and form an HTML code list;

[0031] A text content extraction module, configured to: traverse the HTML code list, extract all text contents in each web page, and convert the long texts of all web pages into short text string lists according to the HTML structure, and retain the text order; wherein, each short text string list belongs to the text list set of the entire web page ethnic group;

[0032] The duplicate removal dictionary filling module is configured to: establish a duplicate removal dictionary, where the keys of the duplicate removal dictionary are text strings and the values are the number of times the strings appear in the entire web page population, traverse the text list set and remove duplicate strings from each short text string list to obtain a duplicate-removed short text string list, and fill the duplicate removal dictionary accordingly;

[0033] The string screening module is configured to: if the value corresponding to the string in the duplicate removal dictionary is greater than the set threshold, the string is removed, otherwise the string is retained;

[0034] The start and end position determination module is configured to: traverse the text list set and locate the start position and end position for each short text string list;

[0035] The body text output module is configured to: select the text from the start position to the end position and output the body text list.

[0036] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the method for extracting the body text of ethnic group web pages based on statistical rules as described in the first aspect of the present invention.

[0037] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for extracting the body text of ethnic group web pages based on statistical rules as described in the first aspect.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] 1. The method and system for extracting the body text of ethnic group web pages based on statistical rules according to the present invention do not require manual participation and special rules, can extract web page texts in different forms, and can retain the basic order structure of the web page text, greatly improving the extraction accuracy and extraction efficiency.

[0040] 2. The method and system for extracting the body text of ethnic group web pages based on statistical rules according to the present invention, compared with the method based on HTML rules, can extract the body text of web pages with non-standard HTML structures. While retaining the web page structure, the extracted content is highly readable and the word order is more coherent, suitable for text processing methods based on deep learning. Description of the Drawings

[0041] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0042] Figure 1Schematic flowchart of the method for extracting ethnic group web page text based on statistical rules provided in Embodiment 1 of the present invention. Detailed implementation mode

[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0044] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0045] It should be noted that the terms used herein are only for describing specific implementation modes and are not intended to limit the exemplary implementation modes of the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0046] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0047] Embodiment 1:

[0048] As Figure 1 shown, Embodiment 1 of the present invention provides a method for extracting ethnic group web page text based on statistical rules, including the following processes:

[0049] Step 1: Define the web page ethnic group.

[0050] Obtain a group of web pages to be processed in the form of a web page ethnic group. The specific grouping method is based on the content of the website navigation bar. For example, all sub-pages on the website list page are divided into a web page ethnic group list.

[0051] Step 2: Extract the original HTML code.

[0052] Traverse the web page ethnic group list obtained in Step 1, and extract the original HTML code of each web page to form an HTML code list.

[0053] Step 3: Extract the text list.

[0054] Traverse the HTML code list obtained in Step 2, extract all text contents in each web page, and convert the long text of the entire web page into a list of short text strings l i ∈L, i ∈ [1,..., n], and retain the text order; L is the set of text lists of the entire web page ethnic group.

[0055] Step 4: Establish a deduplication dictionary.

[0056] Establish a deduplication dictionary D, where the key is the text string and the value is the number of times the string appears in the entire web page population. Traverse L and for each l i Remove duplicate strings to obtain the deduplicated l i ′, and fill the deduplication dictionary D accordingly. The traversed and filled deduplication dictionary D represents the number of times each string appears in the entire web page population.

[0057] Step 5: Set the threshold t.

[0058] If the value corresponding to the string in dictionary D is greater than t, it needs to be excluded; if it is less than t, it needs to be retained.

[0059] Step 6: Traverse L and for each l i Perform Steps 7-9.

[0060] Step 7: Locate the starting position.

[0061] Traverse l from the beginning i and for each string in l i look up the number of occurrences in dictionary D until is found, then the j position is the starting position of the main text.

[0062] Step 8: Locate the ending position.

[0063] Traverse l from the end i and for each string in l i look up the number of occurrences in dictionary D until is found, then the j position is the ending position of the main text.

[0064]

[0065] Step 9: Extract the main text.

[0066] Based on the starting and ending positions of the main text obtained in Steps 7-8, select the text from the starting position to the ending position, and finally output the main text list. If it is necessary to output the main text in text form, special delimiters such as "☆" can be used for separation.

[0066] Example 2:

[0067] Embodiment 2 of the present invention provides a method for extracting the main text of ethnic group web pages based on statistical rules, including the following processes:

[0068] Step 1: Define the web page population.

[0069] Obtain a set of web pages to be processed in the form of a web page group. The specific grouping method is based on the content of the website navigation bar. For example, all sub-pages on the website list page are divided into a web page group list.

[0070] Step 2: Extract the original HTML code.

[0071] Traverse the web page group list obtained in Step 1, extract the original HTML code of each web page, and form an HTML code list.

[0072] Step 3: Extract the text list.

[0073] Traverse the HTML code list obtained in Step 2, extract all the text content in each web page, and according to the structure of HTML, convert the long text of the entire web page into a list of short text strings l i ∈L, i ∈ [1,..., n], and preserve the text order; L is the set of text lists of the entire web page group.

[0074] Step 4: Establish a deduplication dictionary.

[0075] Establish a deduplication dictionary D, where the key is the text string and the value is the number of times the string appears in the entire web page group. Traverse L and for each l i Remove duplicate strings to obtain the deduplicated l′, and fill the deduplication dictionary D accordingly. The traversed and filled deduplication dictionary D represents the number of times each string appears in the entire web page group. i

[0076] Step 5: Set the threshold t.

[0077] If the value corresponding to the string in the dictionary D is greater than t, it needs to be excluded; if it is less than t, it needs to be retained.

[0078] Step 6: Traverse L and for each l i Perform Steps 7 - 9.

[0079] Step 7: Traverse l i Remove the strings whose values corresponding in the dictionary D are greater than the threshold t, and retain the original string order.

[0080] Step 8: Select the position of the first string in the original text of the deduplicated string data group as the starting position of the main text, and select the position of the last string in the original text of the deduplicated string data group as the ending position of the main text.

[0081] Step 9: Extract the main text.

[0082] Based on the starting and ending positions of the main text obtained in steps 7 - 8, select the text from the starting position to the ending position, and finally output a list of main text; if it is necessary to output the main text in text form, special delimiters can be used for splitting, such as "☆".

[0083] Example 3:

[0084] Embodiment 3 of the present invention provides a system for extracting ethnic web page main text based on statistical rules, including:

[0085] A web page ethnic group list generation module, configured to: obtain a group of web pages to be processed in the form of a web page ethnic group, and obtain a web page ethnic group list;

[0086] An HTML code list generation module, configured to: traverse the web page ethnic group list, extract the original HTML code of each web page, and form an HTML code list;

[0087] A text content extraction module, configured to: traverse the HTML code list, extract all text content in each web page, convert the long texts of all web pages into a list of short text strings according to the HTML structure, and retain the text order; among them, each list of short text strings belongs to the text list set of the entire web page ethnic group;

[0088] A duplicate removal dictionary filling module, configured to: establish a duplicate removal dictionary, where the keys of the duplicate removal dictionary are text strings, and the values of the duplicate removal dictionary are the number of times the string appears in the entire web page ethnic group, traverse the text list set and remove duplicate strings from each list of short text strings, obtain a list of short text strings after duplicate removal, and fill the duplicate removal dictionary accordingly;

[0089] A string screening module, configured to: if the value corresponding to the string in the duplicate removal dictionary is greater than a set threshold, the string is excluded, otherwise the string is retained;

[0090] A start and end position determination module, configured to: traverse the text list set and locate the start and end positions for each list of short text strings;

[0091] A main text output module, configured to: select the text from the start position to the end position and output a list of main text.

[0092] The working method of the system is the same as that of the ethnic web page main text extraction method based on statistical rules described in Embodiment 1 or Embodiment 2, and will not be elaborated here.

[0093] Example 4:

[0094] Example 4 of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the steps in the method for extracting ethnic group web page text based on statistical rules described in Example 1 or Example 2 of the present invention are implemented.

[0095] Example 5:

[0096] Example 5 of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for extracting ethnic group web page text based on statistical rules described in Example 1 or Example 2 are implemented.

[0097] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be in the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0098] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the specified functions in Figure 1 one or more flows and / or Figure 1 blocks or multiple blocks.

[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the specified functions in Figure 1One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.

[0101] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0102] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for extracting the main text of ethnic group web pages based on statistical rules, characterized in that: It includes the following processes: Obtain a set of web pages to be processed in the form of a web page ethnic group, and obtain a web page ethnic group list; Traverse the web page ethnic group list, extract the original HTML code of each web page, and form an HTML code list; Traverse the HTML code list, extract all the text content in each web page, and according to the HTML structure, convert the long texts of all web pages into a list of short text strings, and retain the text order; among them, each list of short text strings belongs to the text list set of the entire web page ethnic group; Establish a deduplication dictionary, where the keys of the deduplication dictionary are text strings, and the values of the deduplication dictionary are the number of times the string appears in the entire web page ethnic group. Traverse the text list set and remove duplicate strings from each list of short text strings to obtain a deduplicated list of short text strings, and fill the deduplication dictionary accordingly; If the value corresponding to the string in the deduplication dictionary is greater than the set threshold, the string is excluded, otherwise the string is retained; Traverse the text list set and locate the start position and end position for each list of short text strings; Select the text from the start position to the end position and output the main text list.

2. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: Define the web page ethnic group based on the content of the website navigation bar.

3. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: The deduplication dictionary after traversing and filling represents the number of times each string appears in the entire web page ethnic group.

4. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: Traverse the text list set and locate the start position for each list of short text strings, including: Traverse the list of short text strings from the beginning, and for each string in the list of short text strings Look up the occurrence count in the deduplication dictionary until found Then the position at j is the starting position of the main text, where t is a set threshold value.

5. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: Traverse the text list set and locate the end position for each list of short text strings, including: Traverse the list of short text strings starting from the end, for each string in the list of short text strings Look up the occurrence count in the deduplication dictionary until found Then the position at j is the end position of the main text, where t is a set threshold value.

6. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: Traverse the text list set and locate the start position and end position for each list of short text strings, including: Traverse the list of short text strings, remove the strings whose corresponding values in the deduplication dictionary are greater than the set threshold, and retain the original string order; Select the position of the first string in the original text of the deduplicated string data group as the start position of the main text, and select the position of the last string in the original text of the deduplicated string data group as the end position of the main text.

7. The method for extracting the main text of ethnic group web pages based on statistical rules according to claim 1, characterized in that: When it is necessary to output the main text in text form, use special delimiters for text segmentation.

8. A system for extracting the main text of ethnic group web pages based on statistical rules, characterized in that: It includes: A web page ethnic group list generation module, configured to: obtain a set of web pages to be processed in the form of a web page ethnic group, and obtain a web page ethnic group list; The HTML code list generation module is configured to: traverse the list of web page groups, extract the original HTML code of each web page, and form an HTML code list; The text content extraction module is configured to: traverse the HTML code list, extract all the text content in each web page, convert the long texts of all web pages into a list of short text strings according to the HTML structure, and preserve the text order; among them, each list of short text strings belongs to the text list set of the entire web page group; The duplicate removal dictionary filling module is configured to: establish a duplicate removal dictionary, where the keys of the duplicate removal dictionary are text strings and the values are the number of times the string appears in the entire web page group. Traverse the text list set and remove duplicate strings from each list of short text strings to obtain a list of short text strings after duplicate removal, and fill the duplicate removal dictionary accordingly; The string screening module is configured to: if the value corresponding to the string in the duplicate removal dictionary is greater than the set threshold, the string is excluded; otherwise, the string is retained; The start and end position determination module is configured to: traverse the text list set and locate the start position and end position for each list of short text strings; The main text output module is configured to: select the text from the start position to the end position and output a list of main text.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the method for extracting the main text of ethnic web pages based on statistical rules as described in any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements the steps in the method for extracting the main text of ethnic web pages based on statistical rules as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Web page body text extraction method and apparatus

    CN105183801A

  • Policy webpage text extraction method, system and device and storage medium

    CN111966901A

  • A method and device for extracting a webpage text

    CN109948089A

  • Webpage text recognition processing method and device

    CN110795933A