Multi-feature collaborative full-loop structure webpage text extraction method and system

By integrating web structure and text features, the method enhances the extraction of relevant text from full loop structure web pages, reducing noise and improving accuracy while minimizing development costs.

CN120316280APending Publication Date: 2025-07-15WENZHOU UNIV OUJIANG COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377727.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art requires web page development in full-circular structure web page text extraction, which is not versatile and has high development and maintenance costs, making it difficult to accurately identify and extract content of interest.

Method used

The multi-feature collaboration method is adopted to combine web page structure characteristics and text characteristics to score the tag groups in the web page. The path depth, location, number of punctuation marks, link text ratio, code text ratio, text distribution, similarity and average text length and frequency are used to sort and select the most likely tag group for text extraction.

Benefits of technology

It improves the accuracy and universality of full-loop structure web page text extraction, reduces dependence on web page structure, and reduces development and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005333562480000051
    Figure BDA0005333562480000051
  • Figure BDA0005333562480000052
    Figure BDA0005333562480000052
  • Figure BDA0005333562480000054
    Figure BDA0005333562480000054
Patent Text Reader

Abstract

The invention discloses a multi-feature collaborative full-loop structure webpage text extraction method and system, and the method comprises the following steps: (1) analyzing the structure of a target webpage, taking tags with the same loop structure on the webpage as tag groups, and obtaining all tag groups of the target webpage; (2) carrying out collaborative scoring on all the label groups of the target webpage obtained in the step (1) by adopting various characteristics including webpage structure characteristics and text characteristics to obtain a candidate label group sequence of all the label groups of the target webpage according to the possibility ranking of interested label groups; and (3) extracting text information of the interested tag group. According to the multi-feature collaborative full-loop structure webpage text extraction method provided by the invention, the tag groups in the webpage are analyzed based on the tag path, the score of each tag group is calculated in combination with the webpage structure features and the text features, and then optimization of the target tag group is completed according to the scores, so that text extraction is completed. The method is carried out on the basis of eight weak conditions, so that the method has good universality for different webpages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet technologies, and more specifically, relates to a method and system for extracting web page text with a fully cyclic structure featuring multi-feature collaboration. Background Art

[0002] With the rise of big data, people's demand for big data has become increasingly strong. As an essential basic support for people's work, study, and life, the Internet has various forms of information carrier platforms, such as news portals, forums, blogs, microblogs, WeChat official accounts, e-commerce websites, etc., which store and record a vast amount of data. This provides possibilities for many big data applications. The mining of the vast amount of data existing in these various carriers or platforms can support the macro-decision-making of management departments and the formulation of business strategies of commercial companies, and ultimately be transformed into social or economic benefits. In addition, basic projects such as corpus construction and language large model training also rely on the support of data from the aforementioned various platforms. For this reason, the research on web information extraction has long been a relatively popular research field.

[0003] However, web data is often unstructured, which causes great trouble in data utilization. Fortunately, in web pages, there is a cyclic structure for organizing information, which widely exists in various web pages such as news, forums, blogs, microblogs, WeChat official accounts, etc. Cyclic content refers to the content contained in the cyclic structure in the web page, and the content that is not in the cyclic structure is non-cyclic content. For example, the extraction of titles in web pages and the main text of news portal web pages is generally the extraction of non-cyclic content, while the extraction of user comments on websites, the titles of each post in forum index web pages, and user posts in forum content web pages is the extraction of cyclic content. Almost all the core information in forum pages is in a cyclic hierarchical structure. Since the text content of non-cyclic content web pages is concentrated and the features are more obvious, compared with the content of the cyclic structure, its extraction difficulty is relatively low.

[0004] Among them, when all the core information exists in the cyclic structure, it is called a fully cyclic structure, and a typical representative is a forum; when the core information has a primary and secondary status, and the primary information is independent while the secondary information is in the cyclic structure, it is called a semi-cyclic structure, and typical representatives are news, blogs, microblogs, etc. with comments.

[0005] In structured and semi-structured information extraction, typical examples include forum posts and various user comments with full-loop structures. The data contained in these loop structures often comes from structured databases. In addition, the main texts of some news or blog web pages also have full-loop structures. The loop structure is a relatively common structure in web pages, especially in forum web pages. Because the loop structures of various web pages contain rich and valuable information, which is particularly important for public opinion analysis, commercial marketing, etc., a large number of related studies have been attracted.

[0006] For the extraction method of full-loop structure web page text, generally, web page structure features are used as the main features for text extraction. In addition to using relatively detailed DOM trees and related node attributes to analyze its layout for the loop structure features of web pages, there is also a relatively rough way of parsing, that is, the tag path. Related methods based on tag paths generally only need to parse to the tag name level and generally do not parse the attributes of each tag. However, practical analysis shows that for web pages with full-loop structures, although the target information exists in the full-loop structure, irrelevant information, such as advertisements, signatures, etc., also shows a very regular full-loop structure. Using only the tag paths constructed by DOM nodes easily leads to a large number of noise groups, and even there are too many noise tag paths mixed in the normal groups. Therefore, for the extraction of full-loop structure web page text, it is often necessary to develop and use strict content features for screening according to the web page structure. Its generality is poor, the development cost is high, and the workload of updating and maintenance is large. Summary of the Invention

[0007] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and system for extracting full-loop structure web page text with multi-feature collaboration, aiming to evaluate the text contained in various loop structures on the web page from multiple angles using web page structure features and text features, accurately identify the loop structures containing the text of interest, extract its text, so as to accurately extract the content of interest on the web page, thereby solving the technical problems of the prior art that extracting the text of the full-loop structure of interest requires web page development, poor generality, and high development and maintenance costs.

[0008] To achieve the above object, according to one aspect of the present invention, a method for extracting full-loop structure web page text with multi-feature collaboration is provided, including the following steps:

[0009] (1) Analyze the structure of the target web page, regard the tags with the same loop structure on the web page as a tag group, and obtain all tag groups of the target web page;

[0010] (2) For all tag groups of the target web page obtained in step (1), multiple features including web page structure features and text features are used for collaborative scoring. According to the scores of the collaborative scoring, the tag groups are sorted according to the possibility of being an interested tag group, and a candidate tag group sequence of all tag groups of the target web page sorted according to the possibility of being an interested tag group is obtained;

[0011] (3) According to the candidate tag group sequence in step (2), the candidate tag group with the greatest possibility of being an interested tag group is used as the interested tag group, and the text information of the interested tag group is extracted as the extraction result of the text of the full-cycle structure web page of the target web page.

[0012] Preferably, in the method for extracting the text of the full-cycle structure web page with multi-feature collaboration, in step (1), the tag path method is used to analyze the structure of the target web page, and the tag contents with the same tag path are used as a tag group.

[0013] Preferably, in the method for extracting the text of the full-cycle structure web page with multi-feature collaboration, the web page structure feature in step (2) is the feature of this tag group relative to the web page structure; the text features include the text feature within the element and the text feature between elements. The text feature within the element is the statistical feature of the text content of each element within this tag group, and the text feature between elements is the statistical feature between the text contents of each element. The element refers to the tag content within the tag group, and each element within a tag group has the same tag path.

[0014] Preferably, in the method for extracting the text of the full-cycle structure web page with multi-feature collaboration, the web page structure features include the path depth s i1 and / or the position s i2 ; the path depth is used to represent the depth of the tag group in the web page tag tree structure, denoted as s i1 ; the position is the average position of each element in this tag group in the web page tag tree structure, denoted as where I ij represents the position of the j-th element in the i-th tag group, M represents the total number of tags in the target web page, and n i represents the number of elements in the i-th tag group.

[0015] Preferably, in the method for extracting the text of the full-cycle structure web page with multi-feature collaboration, the text feature within the element includes one or more of the number of punctuation marks s i3 , the ratio of link text s i4 , and the ratio of code text s i5 ;

[0016] The number of punctuation marks is used to represent the total number of punctuation marks in the text within each element in the tag group, where pij Denote the number of punctuation marks in the text content of the j-th element within the label group i; the link text ratio is used to represent the ratio of the total number of links to the total text length within the label group, denoted as where l ij Denote the number of links of the j-th element within the label group i, and t ij Denote the length of the text content of the j-th element within the label group i; the code text ratio is used to represent the ratio of the code length to the text length within the label group, denoted as where c ij Denote the code length of the j-th element within the label group i.

[0017] Preferably, for the multi-feature collaborative full-cycle structure web page text extraction method, the text features between elements include text distribution s i6 , similarity s i7 , and text average length frequency s i8 One or more of them;

[0018] The text distribution is used to represent the balance of text lengths within each element in the label group, denoted as where std() represents calculating the standard deviation of the data; the similarity is used to represent the similarity between the contents of each element in the label group, that is, the average value of the text similarities among the elements in the label group, denoted as where sim(x,y) represents calculating the similarity value of texts x and y; the text average length frequency is used to represent the possibility that the elements of the label group are included in labels at different levels, measured by the occurrence frequency of the average text length of the elements in the label group among the average text lengths of all candidate groups, denoted as where Represents calculating the frequency of the average text length of each label group within the set G composed of all label groups.

[0019] Preferably, for the multi-feature collaborative full-cycle structure web page text extraction method, the multiple features do not include text length.

[0020] Preferably, for the method for extracting web page text with a full cycle structure with multi-feature collaboration, the possibility that the tag group in step (2) is an interested tag group is determined according to the following method: the shallower the depth of the tag group path, the smaller the possibility that it is an interested tag group; the farther the position of the tag group is from the middle of the web page, the smaller the possibility that it is an interested tag group, the smaller the number of punctuation marks in the tag group, the smaller the possibility that it is an interested tag group; the higher the link text ratio of the tag group, the smaller the possibility that it is an interested tag group; the higher the code text ratio of the tag group, the smaller the possibility that it is an interested tag group; the more balanced the text distribution of the tag group, the smaller the possibility that it is an interested tag group; the higher the similarity of the tag group, the smaller the possibility that it is an interested tag group; the lower the frequency of the average length of the text in the tag group, the smaller the possibility that it is an interested tag group.

[0021] Preferably, for the method for extracting web page text with a full cycle structure with multi-feature collaboration, in step (2), the weighted value of the normalized values of the multiple features is used as the possibility score of the tag group being an interested tag group, and the tag groups are sorted according to the size of the score; specifically:

[0022] The possibility score S i of the i-th tag group being an interested tag group is calculated as follows:

[0023]

[0024] where w j is the weight of the j-th feature, and s i ′ j is the normalized value of the j-th feature of the i-th tag group; when the value of S i is larger, it means that the possibility of this tag group being an interested tag group is greater. Among them:

[0025] The normalized value s i1 of the path depth s i is calculated as follows:

[0026]

[0027] where r d is a depth correction factor used to avoid too large a difference in scores for different depths. For example, it can be taken as 10; respectively represent taking the maximum and minimum values of this index among all tag groups.

[0028] The normalized value s i2 of the position s i is calculated as follows:

[0029]

[0030] where, rl is the position factor, whose value ranges from 0 to 1. Generally, it can be taken near the middle of the web page structure. For example, it can be taken as 0.55.

[0031] The number of punctuation marks s i3 The normalized value s of i ′3 is calculated as follows:

[0032]

[0033] where r p is the punctuation correction factor to avoid the case of a zero denominator. It is a positive integer and can be taken as 1.

[0034] The link text ratio s i4 The normalized value s of i ′4 is calculated as follows:

[0035]

[0036] The code text ratio s i5 The normalized value s of i ′5 is calculated as follows:

[0037]

[0038] The text distribution s i6 The normalized value s of i ′6 is calculated as follows:

[0039]

[0040] The similarity s i7 The normalized value s of i ′7 is calculated as follows:

[0041]

[0042] The text average length frequency s i8 The normalized value s of i ′8 is calculated as follows:

[0043] s i ′8 = s i8 .

[0044] Preferably, in the multi - feature collaborative full - cycle structure web page text extraction method, the interested tag group is the tag group with the largest possibility score among the interested tag groups, that is:

[0045] According to another aspect of the present invention, there is provided a multi - feature collaborative full - cycle structure web page text extraction system, which is an electronic device or a non - transient computer - readable storage medium;

[0046] The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the multi-feature collaborative full-cycle structure web page text extraction method provided by the present invention;

[0047] The non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-feature collaborative full-cycle structure web page text extraction method provided by the present invention.

[0048] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0049] The multi-feature collaborative full-cycle structure web page text extraction method provided by the present invention parses out tag groups in a web page based on a tag path, combines web page structure features and text features to judge the possibility that a tag group is an extraction target, and then extracts the most likely target text. It has weak requirements for web page structure information and good versatility for different web pages.

[0050] The present invention calculates the scores of each candidate group using 8 defined feature indicators, and finally completes the selection of the optimal group using the scores. Compared with other methods for text extraction based on tag paths, the innovations in this article are reflected in:

[0051] (1) When performing tag path parsing, only the class attribute of the tag is considered in addition to the tag itself, and other attributes are not considered at all. When all tags have no class attribute, it degrades to a conventional tag path; when there is a class attribute, the class is more conducive to screening out those tags that truly belong to the same group when presented on the page.

[0052] (2) Eight feature indicators and their quantitative calculation methods are proposed. Among them, indicators such as the text distribution, similarity, and text average length frequency within the tag group are innovatively proposed, effectively improving the text extraction effect of the full-cycle structure page. In addition, taking the tag group as the analysis unit is a page-level global perspective, avoiding misjudgment caused by a small number of exceptional tags (such as ultra-short texts, images).

[0053] (3) The selection of the optimal tag group, that is, the tag group of interest, is realized using 8 index features and their weights, completing the extraction of the full-cycle structure web page text, and the selection of the optimal tag group does not require threshold support. Detailed implementation

[0054] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0055] The multi-feature collaborative full-cycle structure web page text extraction method provided by the present invention includes the following steps:

[0056] (1) Analyze the structure of the target web page, take the tags with the same cycle structure on the web page as a tag group, and obtain all tag groups of the target web page; in a preferred solution, use the tag path method to analyze the structure of the target web page, and take the tag contents with the same tag path as a tag group;

[0057] (2) For all tag groups of the target web page obtained in step (1), use multiple features including web page structure features and text features for collaborative scoring, and sort the tag groups according to the scores of the collaborative scoring according to the possibility of being an interesting tag group, and obtain a candidate tag group sequence of all tag groups of the target web page sorted according to the possibility of being an interesting tag group;

[0058] The web page structure feature is the feature of this tag group relative to the web page structure, including the path depth s i1 and / or the position s i2 . The path depth is used to characterize the depth of the tag group in the web page tag tree structure, denoted as s i1 . In a common html web page, that is, the number of tag nodes experienced from the DOM tree root node to any element node in the tag group; generally, for a web page with a cycle structure, the path depth of the text to be extracted is generally relatively large, and of course, it is generally not the maximum value of all tag path depths; the position is the average position of each element in this tag group in the web page tag tree structure, denoted as In a common html web page, where I ij represents the position of the jth element in the tag group i, M represents the total number of tags in the target web page, and n i represents the number of elements in the tag group i; that is, the average value of the tag position serial numbers of each element in the original web page DOM tree. Generally speaking, the position of the text to be extracted tends to be in the middle part of the page.

[0059] The text feature includes the in-element text feature and the inter-element text feature. The in-element text feature is the statistical feature of the text content of each element in this tag group, including the number of punctuation marks s i3 the link text ratio s i4 and the code text ratio si5 one or more of the above; the text feature between elements is the statistical feature between the text contents of each element, including text distribution s i6 similarity s i7 and average text length frequency s i8 one or more of the above. The element refers to the label content within the label group. Each element within a label group has the same label path.

[0060] The number of punctuation marks is used to represent the total number of punctuation marks in the text within each element in the label group, denoted as where p ij represents the number of punctuation marks in the text content of the j-th element in the i-th label group; the link text ratio is used to represent the ratio of the total number of links to the total text length in the label group, denoted as where l ij represents the number of links of the j-th element in the i-th label group, and t ij represents the length of the text content of the j-th element in the i-th label group; the code text ratio is used to represent the ratio of the code length to the text length in the label group, denoted as where c ij represents the code length of the j-th element in the i-th label group.

[0061] The text distribution is used to represent the balance of the text lengths within each element in the label group, denoted as where std() represents calculating the standard deviation of the data. Since in a loop-structured web page, generally speaking, compared with other label groups, the text lengths of the elements within the label group containing the target text are not overly disparate, so they have a relatively small standard deviation. Of course, generally they are not 0 or very small; the similarity is used to represent the similarity between the contents of each element in the label group, that is, the average value of the text similarities within each element in the label group, denoted as where sim(x,y) represents calculating the similarity value between text x and y. For the label group containing the target text, the similarity values of the text within each element are often relatively low; however, for some incorrect candidate label groups, there may be a large amount of repetition. Therefore, this indicator can effectively remove some candidate label groups with such errors; the average text length frequency is used to represent the possibility that the elements of the label group are included in labels at different levels, measured by the occurrence frequency of the average text length of the elements in the label group among the average text lengths of all candidate groups, denoted as where Indicates the frequency of calculating the average text length of each tag group within the set G composed of all tag groups. The original intention of this feature design is that the text to be extracted often exists in tag groups at multiple levels. The root cause is that web designers, based on the flexibility of web page presentation control, often wrap text content through multiple levels of container tags (usually DIV, TABLE, TR, etc.). Therefore, the correct text to be extracted can be obtained from multiple levels, which means that the target text can be extracted from multiple tag groups. Although the depths of these tag groups in the tag tree are slightly different, their average text lengths are the same, and other content rarely has this feature. Therefore, this indicator is also extremely beneficial for the correct recognition and extraction of the target text.

[0062] In addition, it should be noted that many current related methods generally consider text length to be a reliable indicator of the target text, believing that the longer the text length, the more information it provides, and the more likely this text is to be the target text. However, through the analysis of a large number of target texts with full loop structures, it is found that in web pages with full loop structures, the average text length often does not have obvious characteristics. For example, the target text lengths on Weibo are generally short, while those on Zhihu are generally long. Therefore, using text length is extremely likely to lead to misjudgment. For this reason, in order to improve the generality of the text extraction method for full loop structures, it is preferably not to directly use text length as a feature.

[0063] The possibility that a tag group is an interested tag group is determined as follows: The shallower the depth of the tag group path, the smaller the possibility that it is an interested tag group; the farther the position of the tag group is from the middle of the web page, the smaller the possibility that it is an interested tag group; the smaller the number of punctuation marks in the tag group, the smaller the possibility that it is an interested tag group; the higher the link text ratio of the tag group, the smaller the possibility that it is an interested tag group; the higher the code text ratio of the tag group, the smaller the possibility that it is an interested tag group; the more balanced the text distribution of the tag group, the smaller the possibility that it is an interested tag group; the higher the similarity of the tag group, the smaller the possibility that it is an interested tag group; the lower the frequency of the average text length of the tag group appears, the smaller the possibility that it is an interested tag group.

[0064] In a preferred solution, the weighted value of the normalized values of the multiple features is used as the possibility score for a tag group to be an interested tag group, and the tag groups are sorted according to the size of the scores; specifically:

[0065] The possibility score S i for the i-th tag group to be an interested tag group is calculated as follows:

[0066]

[0067] When S iThe larger the value, the greater the likelihood that the tag group is the tag group of interest.

[0068] Among them, w j is the weight of the j-th feature, and s i ′ j is the normalized value of the j-th feature of the i-th tag group; where:

[0069] The normalized value s i1 of the path depth s i ′1 is calculated as follows:

[0070]

[0071] Among them, r d is the depth correction factor, used to avoid too large a difference in scores at different depths. For example, it can be taken as 10; respectively represent the maximum and minimum values of this index among all tag groups.

[0072] The normalized value s i2 of the position s i ′2 is calculated as follows:

[0073]

[0074] Among them, r l is the position factor, whose value ranges from 0 to 1. In general, it can be taken near the middle of the web page structure. For example, it can be taken as 0.55.

[0075] The normalized value s i3 of the number of punctuation marks s i ′3 is calculated as follows:

[0076]

[0077] Among them, r p is the punctuation correction factor, to avoid the case of a zero denominator, and is a positive integer, which can be taken as 1.

[0078] The normalized value s i4 of the link text ratio s i ′4 is calculated as follows:

[0079]

[0080] The normalized value s i5 of the code text ratio s i ′5 is calculated as follows:

[0081]

[0082] The normalized value s i6 of the text distribution s i'6 is calculated as follows:

[0083]

[0084] Similarity s i7 The normalized value s of i '7 is calculated as follows:

[0085]

[0086] Text average length frequency s i8 The normalized value s of i '8 is calculated as follows:

[0087] s i '8 = s i8

[0088] (3) According to the candidate tag group sequence in step (2), take the tag group with the highest possibility of being the tag group of interest as the tag group of interest, and extract the text information of the tag group of interest as the extraction result of the text of the full-cycle structure web page of the target web page.

[0089] In a preferred solution, the tag group of interest is the tag group with the highest possibility score of being the tag group of interest, that is:

[0090] The following are examples:

[0091] The full-cycle structure web page text extraction method with multi-feature collaboration provided in this embodiment includes the following steps:

[0092] (1) Analyze the structure of the target web page, take the tags with the same loop structure on the web page as tag groups, and obtain all tag groups of the target web page; in a preferred solution, use the tag path method to analyze the structure of the target web page, and take the tag contents with the same tag path as a tag group;

[0093] A so-called tag group is a set of HTML codes formed by tag contents with the same tag path in the same web page while maintaining the original order, and each tag content is called an element within the tag group.

[0094] According to the above definition, it can be known that the tag group has the following characteristics.

[0095] 1. Each element in the tag group corresponds to a section of HTML code.

[0096] 2. Each element in the tag group has the same tag path.

[0097] 3. Each element in the tag group is ordered, and its order is sorted from front to back according to the order of the current element in the DOM tree of the current web page.

[0098] In addition, in the vast majority of pages, tags with the same tag path often have similar appearances, which means that they are also very similar in code structure and express similar semantics. In addition, during the DOM tree parsing of the page, each tag is numbered from front to back, and this number is regarded as the position of the tag.

[0099] (2) For all tag groups of the target web page obtained in step (1), multiple features including web page structure features and text features are used for collaborative scoring. According to the scores of the collaborative scoring, the tag groups are sorted according to the possibility of being an interested tag group, and a candidate tag group sequence sorted according to the possibility of being an interested tag group for all tag groups of the target web page is obtained;

[0100] For the convenience of quantitative description below, let there be M tags in the page to be processed, N tag groups (denoted as G), and each tag group is denoted as G i (i = 1, 2,..., N), and the number of elements in the tag group is denoted as n i (i = 1, 2,..., N). For the i-th tag group, the tag positions (numbers) of its respective elements are I ij , the code length is c ij , the number of punctuation marks is p ij , the text is T ij , the text length is t ij , and the number of links is l ij , where the subscript i represents the tag group number and j represents the element sequence number within the group.

[0101] The features adopted in this embodiment are:

[0102] Path depth, that is, the number of tag nodes experienced from the DOM tree root node to any element node within the tag group, denoted as s i1 . Generally, for web pages with a loop structure, the path depth of the tags where the text to be extracted is located is relatively large. Of course, it generally won't be the maximum value of all tag paths either.

[0103] Position, that is, the average value of the tag position serial numbers of each element within the tag group in the original web page DOM tree, denoted as s i2 . Among them, I ij represents the tag position of the j-th element within the current group in the DOM tree. Generally speaking, the position of the text to be extracted is in the middle of the page.

[0104] Number of punctuation marks, that is, the total number of punctuation marks in the text within each element of the tag group, denoted as Generally, there is a positive correlation between the number of punctuation marks and the text length. However, in pages where the overall body content of the page presents a cyclic structure, its cyclic content, such as forum posts, Weibo comments, WeChat official account comments, website comments, e-commerce website product reviews, etc., are often short texts with substandard expressions. Therefore, in fact, they often do not have a strict proportional relationship.

[0105] The link text ratio, that is, the ratio of the total number of links in the tag group to the total text length, is denoted as

[0106] The code text ratio, that is, the ratio of the code length to the text length in the tag group, is denoted as

[0107] The text distribution, that is, the balance of the text lengths within each element in the tag group, is denoted as Among them, std() represents calculating the standard deviation of the specified data. Since in web pages with a cyclic structure, generally speaking, compared with other tag groups, the text lengths of the elements within the tag group containing the target text are not overly disparate. Therefore, they have a relatively small standard deviation, and of course, it generally will not be 0.

[0108] The similarity, that is, the similarity of the text within each element in the tag group. That is, it is the average value of the similarities between the texts within the same tag path group, denoted as Among them, sim(x,y) represents calculating the similarity value between texts x and y. For the tag group containing the target text, the similarity values of the texts within each element are often low; however, for some incorrect candidate tag groups, there may be a large number of repetitions. Therefore, this indicator can effectively remove some candidate tag groups with such errors.

[0109] The frequency of the average text length, that is, the frequency of the average text length of the tag group among the average text lengths of all candidate groups (the possibility of being included in tag groups at different levels), is denoted as Among them represents calculating the frequency of the average text length of each group within all tag groups G. The original intention of designing this indicator is as follows: The text to be extracted often exists in tag groups at multiple levels. The reason is that web designers, for the sake of flexibility in controlling the web page presentation, often wrap the text content with container tags at multiple levels (usually DIV, TABLE, TR, etc.). Therefore, the correct text to be extracted can be obtained from multiple levels. This means that the correct result can be extracted from multiple tag groups. Although these tag groups have slightly different depths in the tag tree, their average text lengths are the same. Other content rarely has this feature, so this indicator is also extremely beneficial for the correct identification and extraction of text.

[0110] The possibility that the tag group is an interested tag group is determined as follows:

[0111] By comprehensively using the above 8 feature indicators and combining the judgment conditions, for each tag group, the indicators are defined as shown in Table 1.

[0112] Table 1 Tag group features and weights

[0113]

[0114] Then, for the i-th tag group, its possibility score of being an interested tag group is as follows:

[0115]

[0116] The optimal tag group can be determined by the following formula. After determining the optimal tag group, the text extraction in the loop structure can be completed.

[0117]

[0118] Based on the tag groups formed after DOM parsing of the web page and the above indicators in this embodiment, we propose a full loop structure web page text extraction method with multi-feature collaboration. The core of this algorithm is as follows:

[0119] Algorithm: Full loop structure web page text extraction

[0120]

[0121] (3) According to the candidate tag group sequence in step (2), take the tag group with the highest possibility of being an interested tag group as the interested tag group, and extract the text information of the interested tag group as the extraction result of the full loop structure web page text of the target web page.

[0122] The interested tag group is the tag group with the highest possibility score of being an interested tag group, that is:

[0123] Performance evaluation of the method:

[0124] In this embodiment, three indicators of accuracy, recall rate, and F1 value are used to evaluate the extraction performance of the method. The evaluation is divided into two levels. First is the extraction evaluation of a single page. Since in the full loop structure web page, the content to be extracted exists dispersedly and may be presented in multiple levels, there may be various situations of errors and omissions in the extraction process. Therefore, it is necessary to evaluate at the level of a single page first. Second is the extraction evaluation at the domain level, that is, using the average value of the extraction performance of all pages under the same domain to evaluate the extraction of the entire domain.

[0125] For a single page, the accuracy rate (P i) Recall rate (R i ) F1 score (F 1i ) are defined as follows:

[0126]

[0127]

[0128] where i is used to identify different pages, and Text extract represents the text of the i-th page extracted by the method of this article, and Text label represents the manually extracted text for the i-th page. lcs(x, y) represents finding the longest common substring of strings x and y.

[0129] For the extraction results of pages under a certain domain, the above three indicators are defined as follows:

[0130]

[0131] where n represents the size of the dataset in the domain to be evaluated.

[0132] Test data and basic parameters:

[0133] The experimental data in this embodiment is mainly collected randomly from the network, that is, randomly collected from a specific domain name, and finally randomly selected and merged from the randomly collected pages into an experimental dataset. The amount of data in each domain in the dataset is the same. Finally, there are 29 domains and 2900 pages in total. For specific domains, see Table 2.

[0134] The relevant parameter settings for the experiment are as follows:

[0135] (1) The weights of 8 feature indicators are: w1 = 0.04, w2 = 0.08, w3 = 0.2, w4 = 0.23, w5 = 0.16, w6 = 0.13, w7 = 0.12, w8 = 0.04.

[0136] (2) Depth correction factor r d = 10.

[0137] (3) Position factor r l = 0.5532.

[0138] (4) Punctuation correction factor r p = 1.

[0139] The extraction results of the experiment are shown in Table 2 below.

[0140] Table 2 Extraction experiment results

[0141]

[0142] As can be seen from the above table, the average accuracy rate of this embodiment reaches 98%. In comparison, the average recall rate is slightly lower, at 92%, and the F1 value reaches 95%. That is to say, among the texts extracted in this embodiment, 98% are correct, and only 2% of the extracted texts are incorrect; from the perspective of the texts extracted manually, 92% of the texts are correctly extracted, and only 8% are not successfully extracted. From the F1 value, the method in this paper generally performs well and can better meet the actual needs.

[0143] This embodiment proposes a method for collaborative extraction of multiple features for the problem of extracting text content in web pages with a loop structure. This method first makes full use of the locally identical structural features in the web page to construct a tag group; secondly, it defines 8 feature indicators and corresponding weak conditions for text extraction, calculates the comprehensive scores of each candidate tag group through these 8 feature indicators, and finally uses the scores of each group for sorting to realize the optimization of the extraction results. When optimizing each candidate tag group, this method has a global perspective; it collaboratively uses 8 feature indicators, and the corresponding hypothetical conditions of each indicator are weak conditions, with a wide range of applications; in addition, this method also has the characteristic of being language-independent; and it does not require threshold support. These determine that the method in this paper has good generality in loop-type web pages.

[0144] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for extracting web page text with a full-cycle structure featuring multi-feature collaboration, characterized in that It includes the following steps: (1) Analyze the structure of the target web page, take the tags with the same loop structure on the web page as a tag group, and obtain all tag groups of the target web page; (2) For all tag groups of the target web page obtained in step (1), use multiple features including web page structure features and text features for collaborative scoring, and sort the tag groups according to the possibility of being an interested tag group based on the scores of the collaborative scoring, to obtain a candidate tag group sequence of all tag groups of the target web page sorted according to the possibility of being an interested tag group; (3) According to the candidate tag group sequence in step (2), take the candidate tag group with the greatest possibility of being an interested tag group as the interested tag group, and extract the text information of the interested tag group as the extraction result of the full loop structure web page text of the target web page.

2. The method for extracting web page text with a full-cycle structure of multi-feature collaboration according to claim 1, characterized in that, In step (1), the tag path method is used to analyze the structure of the target web page, and the tag content with the same tag path is taken as a tag group.

3. The method for extracting web page text with a full-cycle structure featuring multi-feature collaboration as claimed in claim 1, wherein The web page structure feature in step (2) is the feature of this tag group relative to the web page structure; the text feature includes the in - element text feature and the inter - element text feature. The in - element text feature is the statistical feature of the text content of each element in this tag group, and the inter - element text feature is the statistical feature between the text contents of each element; the element refers to the tag content in the tag group, and each element in a tag group has the same tag path.

4. The method for extracting web page text with a full-cycle structure of multi-feature collaboration as claimed in claim 3, wherein The web page structure features include the path depth s i1 and / or the position s i2 ; the path depth is used to characterize the depth of the tag group in the web page tag tree structure, denoted as s i1 ; the position is the average position of each element in the tag group in the web page tag tree structure, denoted as where I ij represents the position of the j-th element in the tag group i, M represents the total number of tags in the target web page, and n i represents the number of elements in the tag group i.

5. The method for extracting web page text with a full-cycle structure featuring multi-feature collaboration according to claim 3, characterized in that, The text features within the element include the number of punctuation marks s i3 , the link text ratio s i4 , and the code text ratio s i5 , or one or more of them; The number of punctuation marks is used to represent the total number of punctuation marks in the text within each element in the tag group. where p ij represents the number of punctuation marks in the text content of the j-th element in the i-th tag group; the link text ratio is used to represent the ratio of the total number of links to the total text length in the tag group, denoted as where l ij represents the number of links of the j-th element in the i-th tag group, and t ij represents the length of the text content of the j-th element in the i-th tag group; the code text ratio is used to represent the ratio of the code length to the text length in the tag group, denoted as where c ij represents the code length of the j-th element in the i-th tag group.

6. The method for extracting web page text with a full-cycle structure featuring multi-feature collaboration as claimed in claim 3, wherein The text features between the elements include text distribution s i6 , similarity s i7 , and text average length frequency s i8 ; one or more of them The text distribution is used to characterize the balance of the text lengths within each element in the tag group, denoted as where std() represents calculating the standard deviation of the data; the similarity is used to characterize the similarity between the contents of each element in the tag group, that is, the average value of the text similarities in each element in the tag group, denoted as where sim(x, y) represents calculating the similarity value between texts x and y; the text average length frequency is used to characterize the possibility that the elements of the tag group are included in tags at different levels, measured by the occurrence frequency of the average text length of the elements in the tag group among the average text lengths of all candidate groups, denoted as where represents calculating the frequency of the average text length of each tag group within the set G composed of all tag groups.

7. The method for extracting web page text with a full-cycle structure featuring multi-feature collaboration according to claim 1, wherein The possibility that a tag group in step (2) is an interested tag group is determined as follows: the shallower the tag group path depth, the smaller the possibility that it is an interested tag group; the farther the tag group position is from the middle of the web page, the smaller the possibility that it is an interested tag group; the smaller the number of punctuation marks in the tag group, the smaller the possibility that it is an interested tag group; the higher the link text ratio of the tag group, the smaller the possibility that it is an interested tag group; the higher the code text ratio of the tag group, the smaller the possibility that it is an interested tag group; the more balanced the text distribution of the tag group, the smaller the possibility that it is an interested tag group; the higher the similarity of the tag group, the smaller the possibility that it is an interested tag group; the lower the frequency of the average text length of the tag group appears, the smaller the possibility that it is an interested tag group.

8. The method for extracting web page text with a full-cycle structure of multi-feature collaboration according to claim 1, characterized in that, In step (2), the weighted value of the normalized values of the multiple features is used as the possibility score that the tag group is an interested tag group, and the tag groups are sorted according to the size of the scores; specifically: The probability that the i-th tag group is the tag group of interest is measured by S i and is calculated as follows: where, w j is the weight of the j-th feature, and s′ ij is the normalized value of the j-th feature of the i-th label group; when the S i value is larger, it means that the possibility of this label group being an interested label group is greater; where: Path depth s i1 Normalized value s' i1 Is calculated as follows: where r d is the depth correction factor; respectively represent the maximum and minimum values of this indicator in all tag groups. Position s i2 The normalized value s' i2 is calculated as follows: where r l is the position factor, which can generally be taken near the middle of the web page structure; The number of punctuation marks s i3 The normalized value s' i3 is calculated as follows: where r p is a punctuation correction factor to avoid the case of a zero denominator, which is a positive integer and can be taken as 1. The link text ratio s i4 The normalized value s' i4 is calculated as follows: The code text ratio s i5 The normalized value s' i5 is calculated as follows: Text distribution s i6 The normalized value s' i6 is calculated as follows: Similarity s i7 The normalized value s' i7 is calculated as follows: Text average length frequency s i8 Normalized value s' i8 The calculation is as follows: s′ i8 = s i8 .

9. The method for extracting web page text with a full-cycle structure featuring multi-feature collaboration according to any one of claims 1 to 8, characterized in that, The group of tags of interest is the group of tags with the highest likelihood score for being the group of tags of interest, i.e.:

10. A full-cycle structure web text extraction system with multi-feature collaboration, characterized in that It is an electronic device or a non - transitory computer - readable storage medium; The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the multi - feature collaborative full loop structure web page text extraction method as described in any one of claims 1 to 9; The non - transitory computer - readable storage medium has a computer program stored thereon. When the computer program is executed by the processor, it implements the steps of the multi - feature collaborative full loop structure web page text extraction method as described in any one of claims 1 to 9.