A method and system for extracting the main text of news web pages based on semantic similarity
By extracting news titles and related paragraphs based on semantic similarity, defining the text scope, solving the problems of high error rate and missing paragraphs in the prior art, and achieving more accurate and complete web page text extraction.
Patent Information
- Application Number
- CN202510278653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing web page text extraction methods are insufficient in determining thresholds, clustering-based methods are susceptible to noise, and text density-based methods have high requirements for typesetting, resulting in high extraction error rate or missing text paragraphs.
Based on semantic similarity, news titles and most relevant paragraphs are extracted, and the common prefix length and semantic similarity of multi-class label text content are calculated, the text scope is defined and the web page text is extracted.
It improves the accuracy and completeness of text extraction, reduces false extraction and omissions, and enhances the accuracy of extraction results.
Smart Images

Figure CN119783682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more particularly, to a method and system for extracting the main text of news web pages based on semantic similarity. Background Art
[0002] At present, the methods for extracting the main text of web pages mainly include the following categories:
[0003] a) DOM tree-based method: By building a DOM tree and traversing the tree, various web page noises are identified and removed, and the remaining content is used as the extraction result.
[0004] b) Tag density-based method: Utilize the difference in HTML tag density to distinguish the main text and noises.
[0005] c) Machine learning-based method: Linearly reconstruct the web page source code, and obtain the context paragraphs of the web page through text classification and clustering.
[0006] d) Text density-based method: This method first removes the HTML tags of the web page, leaving only the pure text content, then constructs a text density function of the web page, and sets a threshold based on this. The lines or blocks with text density higher than the threshold are regarded as the main text part.
[0007] In the tag density-based method, the accuracy of main text extraction highly depends on the determination of the threshold, and the error rate is relatively high in actual use. The clustering-based method is easily affected by noise data, often including irrelevant texts or missing main text paragraphs. The text density-based method has high requirements for web page layout, that is, this method assumes that the text density of the main text is the highest throughout the web page, but this assumption is too idealistic. For example, for web pages containing a large amount of comment content, this method often extracts the comments as the main text. Summary of the Invention
[0008] In view of the above problems, the present invention proposes a method for extracting the main text of news web pages based on semantic similarity, including:
[0009] For a target news web page, extract the news title of the target news web page based on semantic similarity, and screen out the most relevant paragraphs of the news title;
[0010] Based on the most relevant paragraphs, define the main text range corresponding to the news title;
[0011] Extract the main text of the news web page within the main text range;
[0012] The step of extracting the news title of the target news web page based on semantic similarity for the target news web page includes:
[0013] Extract multiple types of tags from the target news web page and determine the common prefix length of the text content of each type of tag;
[0014] If the common prefix length of the text content of at least two types of tags is greater than 3, then extract the common prefix of the text content with a common prefix length greater than 3 as the news title;
[0015] If the common prefix length of the text content of at least two types of tags does not reach more than 3, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, filter out the text of a preset length as the target text, perform vectorization processing on the target text, and calculate the semantic similarity pairwise for the vectorized target text, and select the target text with the largest sum of similarity scores as the news title.
[0016] Optionally, multiple types of tags include:
[0017] title tag, h1 tag, and h2 tag.
[0018] Optionally, the preset length is less than 30 characters.
[0019] Optionally, screening out the most relevant paragraphs of the news title includes:
[0020] Based on the news title, take the leaf nodes of the web page document tree as potential text paragraphs, and calculate the semantic similarity between the news title and the potential text paragraphs in turn. Based on the semantic similarity, screen out the most relevant paragraphs.
[0021] Optionally, the most relevant paragraphs are the top 5 paragraphs with the highest semantic similarity between the potential text paragraphs and the news title.
[0022] Optionally, based on the most relevant paragraphs, defining the body range corresponding to the news title includes:
[0023] Determine the potential paragraph nodes of the most relevant paragraphs, and on the web page document tree, find the nearest common ancestor node to the potential paragraph nodes, and use the common ancestor node as the potential news body root tag. After all searches are completed, count the potential news body root tags, and use the potential news body root tag that appears the most times as the body root tag, and define the body range with the body root tag.
[0024] Optionally, extracting the news web page body within the body range includes:
[0025] Traverse the child nodes of the body root tag on the web page document tree in turn, extract and parse the text or picture content in the child nodes, and splice the parsed text or picture content in order to obtain the news web page body.
[0026] On the other hand, the present invention also proposes a news web page body extraction system based on semantic similarity, including:
[0027] A news title extraction unit, configured to extract the news title of the target news web page based on semantic similarity for the target news web page, and screen out the most relevant paragraph of the news title;
[0028] A body definition unit, configured to define the body range corresponding to the news title based on the most relevant paragraph;
[0029] A body extraction unit, configured to extract the news web page body within the body range;
[0030] The extraction of the news title of the target news web page based on semantic similarity includes:
[0031] Extract multiple types of tags of the target news web page, and determine the common prefix length of the text content of each type of tag;
[0032] If the common prefix length of the text content of at least two types of tags is greater than 3, then extract the common prefix of the text content with a common prefix length greater than 3 as the news title;
[0033] If the common prefix length of the text content of at least two types of tags does not reach more than 3, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, screen out the text of a preset length as the target text, perform vectorization processing on the target text, and calculate the semantic similarity pairwise for the vectorized target text, and select the target text with the largest sum of similarity scores as the news title.
[0034] On the other hand, the present invention also provides a computing device, including: one or more processors;
[0035] The processor is configured to execute one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the method as described above is implemented.
[0037] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the method as described above is implemented.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] The present invention provides a method for extracting the main text of a news web page based on semantic similarity, including: for a target news web page, extracting the news title of the target news web page based on semantic similarity, and screening out the most relevant paragraphs of the news title; based on the most relevant paragraphs, defining the main text range corresponding to the news title; and within the main text range, extracting the main text of the news web page. The present invention uses the title as a clue and text semantics as a search tool to define the main text range, which not only strengthens the correctness of main text extraction but also ensures the integrity. Brief Description of the Drawings
[0040] Figure 1 It is a flowchart of the method of the present invention;
[0041] Figure 2 It is a schematic diagram of the method of the present invention;
[0042] Figure 3 It is a structural diagram of the system of the present invention. Detailed Embodiments
[0043] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and their effects according to the present invention.
[0044] Embodiment 1:
[0045] The present invention proposes a method for extracting the main text of a news web page based on semantic similarity, as Figure 1 shown, including:
[0046] Step 1: For a target news web page, extract the news title of the target news web page based on semantic similarity, and screen out the most relevant paragraphs of the news title;
[0047] Step 2: Based on the most relevant paragraphs, define the main text range corresponding to the news title;
[0048] Step 3: Within the main text range, extract the main text of the news web page.
[0049] Among them, for a target news web page, extracting the news title of the target news web page based on semantic similarity includes:
[0050] Extracting multiple types of tags of the target news web page, and determining the common prefix length of the text content of each type of tag;
[0051] If the common prefix length of the text content of at least two types of tags is greater than 3, then use the common prefix of the text content with a common prefix length greater than 3 as the news title for extraction;
[0052] If the common prefix length of at least two types of tag text content does not exceed 3, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, filter out the text of a preset length as the target text. Vectorize the target text, and calculate the semantic similarity pairwise for the vectorized target text. Select the target text with the largest sum of similarity scores as the news title.
[0053] Among them, multiple types of tags include:
[0054] title tag, h1 tag, and h2 tag.
[0055] Among them, the preset length is less than 30 characters.
[0056] Among them, screening out the most relevant paragraphs of the news title includes:
[0057] Based on the news title, take the leaf nodes of the web page document tree as potential text paragraphs, and calculate the semantic similarity between the news title and the potential text paragraphs in turn. Based on the semantic similarity, screen out the most relevant paragraphs.
[0058] Among them, the most relevant paragraphs are the top 5 paragraphs with the highest semantic similarity between the potential text paragraphs and the news title.
[0059] Among them, defining the body range corresponding to the news title based on the most relevant paragraphs includes:
[0060] Determine the potential paragraph nodes of the most relevant paragraphs, and on the web page document tree, find the nearest common ancestor node to the potential paragraph nodes, and take the common ancestor node as the potential news body root tag. After all searches are completed, count the potential news body root tags, and take the potential news body root tag that appears the most times as the body root tag, and define the body range with the body root tag.
[0061] Among them, extracting the news web page body within the body range includes:
[0062] Traverse the child nodes of the body root tag on the web page document tree in turn, extract and parse the text or picture content in the child nodes, and splice the parsed text or picture content in order to obtain the news web page body.
[0063] The present invention will be further described below in conjunction with specific cases:
[0064] The principle of the specific case process is as Figure 2 shown, mainly involving stages such as news title recognition, topK paragraph tag positioning, body range definition, and body extraction. The specific implementation solutions for each stage are as follows:
[0065] News title recognition:
[0066] When the common prefix length of the text content of at least two of the title tag, h1 tag, and h2 tag is greater than 3, the common prefix is determined as the news title. When the above conditions are not met, extract the leaf nodes of the web page document tree, filter the text with a length of less than 30 characters, vectorize it, calculate the semantic similarity pairwise, and select the one with the largest sum of similarity scores as the news title.
[0067] topK paragraph tag positioning:
[0068] Using the obtained title, take the leaf nodes in the document tree as potential text paragraphs, calculate the semantic similarity between the title and the potential text paragraphs in turn, select the top 5 most relevant paragraphs, and record their node positions.
[0069] Body scope definition:
[0070] For the found top 5 potential paragraph nodes, on the document tree, find their nearest common ancestor nodes pairwise as the potential news body root tags. After all the searches are completed, count the potential body root tags, and use the root tag with the most occurrences as the final body root tag.
[0071] Body extraction:
[0072] Traverse the child nodes of the body tag on the document tree in turn, extract and parse the text or picture content in the child nodes, and finally splice the content extracted by the traversal in order to obtain the final news body.
[0073] Embodiment 2:
[0074] The present invention also proposes a news web page body extraction system 200 based on semantic similarity, as Figure 3 shown, including:
[0075] A news title extraction unit 201, configured to extract the news title of the target news web page based on semantic similarity for the target news web page, and screen out the most relevant paragraphs of the news title;
[0076] A body definition unit 202, configured to define the body range corresponding to the news title based on the most relevant paragraphs;
[0077] A body extraction unit 203, configured to extract the news web page body within the body range;
[0078] The extracting the news title of the target news web page based on semantic similarity includes:
[0079] Extract multiple types of tags from the target news web page and determine the common prefix length of the text content of each type of tag;
[0080] If the common prefix length of the text content of at least two types of tags is greater than 3, extract the common prefix of the text content with a common prefix length greater than 3 as the news title;
[0081] If the condition that the common prefix length of the text content of at least two types of tags is greater than 3 is not met, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, filter out text of a preset length as the target text, perform vectorization processing on the target text, and calculate the semantic similarity pairwise for the vectorized target text, and select the target text with the largest sum of similarity scores as the news title.
[0082] The present invention uses the title as a clue and text semantics as a search tool, and then defines the scope of the text body, which not only strengthens the correctness of text body extraction but also ensures the integrity.
[0083] Embodiment 3:
[0084] Based on the same inventive concept, the present invention also provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of the method in the above embodiments.
[0085] Embodiment 4:
[0086] Based on the same inventive concept, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by the processor. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of the method in the above embodiments.
[0087] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above in a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to form equivalent embodiments with equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for extracting the main text of news web pages based on semantic similarity, characterized in that, Including: For a target news web page, extract the news title of the target news web page based on semantic similarity, and screen out the most relevant paragraphs of the news title; Based on the most relevant paragraphs, define the body text range corresponding to the news title; Within the body text range, extract the news web page body text; The extracting the news title of the target news web page based on semantic similarity for the target news web page includes: Extract multiple types of tags of the target news web page, and determine the common prefix length of the text content of each type of tag; If the common prefix lengths of the text contents of at least two types of tags are greater than 3, then use the common prefix of the text contents with a common prefix length greater than 3 as the news title for extraction; If the condition that the common prefix lengths of at least two types of tags are greater than 3 is not met, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, screen out text of a preset length as the target text, perform vectorization processing on the target text, and calculate the semantic similarity pairwise for the vectorized target text, and select the target text with the largest sum of similarity scores as the news title; The screening out the most relevant paragraphs of the news title includes: Based on the news title, use the leaf nodes of the web page document tree as potential text paragraphs, calculate the semantic similarity between the news title and the potential text paragraphs in sequence, and screen out the most relevant paragraphs based on the semantic similarity; The defining the body text range corresponding to the news title based on the most relevant paragraphs includes: Determine the potential paragraph nodes of the most relevant paragraphs, search on the web page document tree for the nearest common ancestor node to the potential paragraph nodes, and use the common ancestor node as the potential news body root tag. After all searches are completed, count the potential news body root tags, and use the potential news body root tag that appears the most times as the body text root tag, and define the body text range with the body text root tag; The extracting the news web page body text within the body text range includes: Traverse the child nodes of the body text root tag on the web page document tree in sequence, extract and parse the text or picture content in the child nodes, and splice the parsed text or picture content in order to obtain the news web page body text.
2. The method for extracting the main text of a news web page according to claim 1, wherein The multiple types of tags include: title tag, h1 tag, and h2 tag.
3. The method for extracting the main text of a news web page according to claim 1, wherein, The preset length is less than 30 characters.
4. The method for extracting the main text of a news web page according to claim 1, wherein, The most relevant paragraphs are the top 5 paragraphs with the highest semantic similarity between the potential text paragraphs and the news title.
5. A news web page body extraction system based on semantic similarity, characterized in that, Including: A news title extraction unit for extracting the news title of a target news web page based on semantic similarity for the target news web page, and screening out the most relevant paragraphs of the news title; A body text definition unit for defining the body text range corresponding to the news title based on the most relevant paragraphs; A body text extraction unit for extracting the news web page body text within the body text range; The extracting the news title of the target news web page based on semantic similarity for the target news web page includes: Extract multiple types of tags of the target news web page, and determine the common prefix length of the text content of each type of tag; If the common prefix length of the text content of at least two types of tag texts is greater than 3, then extract the common prefix of the text content with a common prefix length greater than 3 as the news title; If the common prefix length of the text content of at least two types of tag texts does not reach greater than 3, determine the leaf nodes of the web page document tree of the target news web page, and based on the leaf nodes of the web page document tree, filter out the text of a preset length as the target text, perform vectorization processing on the target text, and calculate the semantic similarity pairwise for the vectorized target text, and select the target text with the largest sum of similarity scores as the news title; The most relevant paragraphs for screening out the news title include: Based on the news title, take the leaf nodes of the web page document tree as potential text paragraphs, and calculate the semantic similarity between the news title and the potential text paragraphs in sequence, and filter out the most relevant paragraphs based on the semantic similarity; Defining the body range corresponding to the news title based on the most relevant paragraphs includes: Determine the potential paragraph nodes of the most relevant paragraphs, and on the web page document tree, find the nearest common ancestor node to the potential paragraph nodes, and take the common ancestor node as the potential news body root tag. After all searches are completed, count the potential news body root tags, and take the potential news body root tag that appears the most times as the body root tag, and define the body range with the body root tag; Extracting the news web page body within the body range includes: Traverse the child nodes of the body root tag on the web page document tree in sequence, extract and parse the text or picture content in the child nodes, and splice the parsed text or picture content in order to obtain the news web page body.
6. A computer device, characterized in that, Includes: One or more processors; The processor is used to execute one or more programs; When the one or more programs are executed by the one or more processors, the method described in any one of claims 1-4 is implemented.
7. A computer-readable storage medium, characterized in that, There is a computer program stored thereon, and when the computer program is executed, the method described in any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Automatic news webpage element extracting method
CN102750390A
News key information extraction method and system
CN106021392A