Web page body extraction method, device, equipment and storage medium

Through feature extraction, encoding and recall processing, the web page type is determined, and the regular rules and web page text extraction model is combined, the problems of low efficiency and low accuracy of traditional methods are solved, and efficient and accurate web page text extraction is achieved.

CN115344772BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210989853.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-05-27
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Traditional web page text extraction methods are inefficient and have low accuracy. Template-based algorithms need to be updated frequently, while statistics-based algorithms are prone to extract irrelevant content.

Method used

By obtaining the features of the web page to be extracted, encoded as a set of vectors, and determining the web page type through recall processing and classification tag analysis. Depending on the type, use regular rules or trained web text extraction model to extract web text.

Benefits of technology

It improves the efficiency and accuracy of web page body extraction, avoids the extraction of irrelevant content, and adapts to changes in different web page types and patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344772B_ABST
    Figure CN115344772B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent decision-making, and discloses a method for extracting web page text, including: extracting features from a web page to be extracted to obtain a web page data feature set, and encoding the web page data feature set to obtain a web page data vector set; recalling the web page data vector set to obtain an index web page data set, and determining the web page type corresponding to the web page to be extracted by analyzing the classification label to which the index web page data set belongs; judging whether the web page type is a text-based web page; when the web page type is not a text-based web page, extracting the web page text according to regular rules; when the web page type is a text-based web page, extracting the web page text using a web page text extraction model. The present invention also relates to a blockchain technology, and the web page text can be stored in a blockchain node. The present invention also proposes a device, equipment, and medium based on web page text extraction. The present invention can improve the efficiency and accuracy of web page text extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making, and particularly to a method, device, equipment and storage medium for extracting web page text. Background Art

[0002] Web page text extraction refers to filtering out some web page noises during the process of browsing a web page and only extracting the content of the web page text. For example, some financial product web pages usually include: web page theme, web page text content, advertisement information, external links and navigation bars, etc. Except for the web page theme and web page text content, the rest of the web page related information can be regarded as web page noises.

[0003] Traditional web page text extraction methods are generally algorithms based on web page templates and algorithms based on statistics. However, these traditional methods have two problems. On the one hand, since the algorithm based on templates needs to rewrite the wrapper when different web page patterns or web page structures change, the efficiency of web page text extraction is low; on the other hand, since the algorithm based on statistics calculates the corresponding web page text density and link density by counting the number of words, the number of links, the number of tag characters, etc. on the web page, and determines the content of the web page text according to the text density and link density, some content unrelated to the web page text is often extracted during the web page text extraction process, resulting in low accuracy of web page text extraction. Summary of the Invention

[0004] The present invention provides a method, device, equipment and storage medium for extracting web page text, and its main purpose is to improve the efficiency and accuracy of web page text extraction.

[0005] To achieve the above object, the present invention provides a method for extracting web page text, including:

[0006] Obtain a web page to be extracted, perform feature extraction on the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set;

[0007] Perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification label to which the indexed web page data set belongs;

[0008] Judge whether the web page type of the web page to be extracted is a text-based web page;

[0009] When the web page type of the web page to be extracted is not a text-based web page, extract the web page text of the web page to be extracted according to a preset regular rule;

[0010] When the web page type of the web page to be extracted is a text-based web page, use the trained web page body extraction model to extract the web page body of the web page to be extracted.

[0011] Optionally, the recall process on the web page data vector set to obtain the indexed web page data set includes:

[0012] Obtain the vector labels of the web page data vector set, and create partition regions according to the vector labels using a preset open-source vector database;

[0013] Store the web page data vector set into the partition regions, and create indexes for the web page data vector sets in each partition region to obtain the indexed web page data set.

[0014] Optionally, determining the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs includes:

[0015] Obtain the web page data of the web page to be extracted, and select the indexed web page data most similar to the web page data from the indexed web page data set as the pre-classified web page label;

[0016] Select the web page label with the most occurrences in the pre-classified web page labels as the web page type corresponding to the web page to be extracted.

[0017] Optionally, using the trained web page body extraction model to extract the web page body of the web page to be extracted includes:

[0018] Use the bidirectional long short-term memory network in the trained web page body extraction model to encode the web page to be extracted to obtain an encoded data set;

[0019] Use the unidirectional long short-term memory network in the web page body extraction model to decode the encoded data set to obtain a decoded data set;

[0020] Input the decoded data set into a preset activation function to obtain an activation probability value, and obtain the web page body according to the activation probability value.

[0021] Optionally, using the bidirectional long short-term memory network in the trained web page body extraction model to encode the web page to be extracted to obtain an encoded data set includes:

[0022] Use the input gate in the bidirectional long short-term memory network to calculate the state value of the web page to be extracted;

[0023] Use the forget gate in the bidirectional long short-term memory network to calculate the activation value of the web page to be extracted;

[0024] Calculate the status update value of the web page to be extracted according to the status value and the activation value;

[0025] Use the output gate in the bidirectional long short-term memory network to calculate the encoded data set corresponding to the status update value.

[0026] Optionally, the extracting the web page body of the web page to be extracted according to a preset regular rule includes:

[0027] Obtain the web page source code of the web page to be extracted, and determine the position of the web page body in the web page according to the web page source code;

[0028] When it is recognized that the web page to be extracted is an image-type web page, use a preset image regular rule to extract the text from the web page body position to obtain the web page body;

[0029] When it is recognized that the web page to be extracted is a link-type web page, use a preset link regular rule to extract the web page link from the web page body position to obtain the web page body.

[0030] Optionally, the extracting features from the web page to be extracted to obtain a web page data feature set includes:

[0031] Convert the web page to be extracted into a text web page, perform word segmentation on the text web page to obtain a word segmentation text set;

[0032] Use a preset algorithm to calculate the weight of each word in the word segmentation text set to obtain word weights;

[0033] Extract the words with word weights greater than a preset threshold from the word segmentation text set as web page keywords;

[0034] Perform part-of-speech tagging on the web page keywords according to a preset dictionary to determine the part-of-speech of the web page keywords;

[0035] Determine the web page data feature set of the web page to be extracted according to the part-of-speech of the web page keywords.

[0036] To solve the above problems, the present invention also provides a web page body extraction device, and the device includes:

[0037] A web page feature extraction module, configured to obtain a web page to be extracted, extract features from the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set;

[0038] A web page type recognition module, configured to perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification label to which the indexed web page data set belongs;

[0039] A web page body extraction module, configured to determine whether the web page type of the to-be-extracted web page is a text-based web page; when the web page type of the to-be-extracted web page is not a text-based web page, extract the web page body of the to-be-extracted web page according to a preset regular rule; when the web page type of the to-be-extracted web page is a text-based web page, extract the web page body of the to-be-extracted web page by using a trained web page body extraction model.

[0040] To solve the above problems, the present invention further provides an electronic device, which includes:

[0041] A memory for storing at least one computer program; and

[0042] A processor for executing the computer program stored in the memory to implement the above-mentioned web page body extraction method.

[0043] To solve the above problems, the present invention further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned web page body extraction method.

[0044] In an embodiment of the present invention, first, by performing feature extraction on the to-be-extracted web page, a web page data feature set is obtained, and the web page data feature set is encoded to obtain a web page data vector set, so that the main feature data in the to-be-extracted web page can be extracted and some useless words can be removed, which is convenient for improving the efficiency of subsequent web page body extraction; secondly, by performing a recall process on the web page data vector set, an indexed web page data set is obtained, an index can be created for each web page data vector, and by analyzing the classification labels to which the indexed web page data set belongs, the web page type corresponding to the to-be-extracted web page can be accurately identified, which is convenient for applying different methods for extraction to different web page types subsequently; finally, when it is identified that the web page type is not a text-based web page, the web page body is extracted by using a regular rule, and when it is identified that the web page type is a text-based web page, the web page body of the to-be-extracted web page is extracted by using a trained web page body extraction model, which can avoid extracting some content irrelevant to the web page body during the extraction process, improve the accuracy of web page body extraction, and different web page types can apply different methods for targeted extraction of the web page body, and there is no need to rewrite the wrapper when different web page modes or web page structures change, which improves the efficiency of web page body extraction. Therefore, the web page body extraction method, device, device and storage medium proposed in the embodiment of the present invention can improve the efficiency and accuracy of web page body extraction. Description of the Drawings

[0045] Figure 1Schematic flowchart of a method for extracting web page text according to an embodiment of the present invention;

[0046] Figure 2 Schematic detailed flowchart of a step in the method for extracting web page text according to an embodiment of the present invention;

[0047] Figure 3 Schematic detailed flowchart of a step in the method for extracting web page text according to an embodiment of the present invention;

[0048] Figure 4 Schematic module diagram of a device for extracting web page text according to an embodiment of the present invention;

[0049] Figure 5 Schematic internal structure diagram of an electronic device for implementing the method for extracting web page text according to an embodiment of the present invention;

[0050] The realization, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0051] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0052] An embodiment of the present invention provides a method for extracting web page text. The execution subject of the method for extracting web page text includes but is not limited to at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the method for extracting web page text can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0053] Referring to Figure 1 the schematic flowchart of the method for extracting web page text provided in an embodiment of the present invention shown in the figure, in the embodiment of the present invention, the method for extracting web page text includes the following steps S1 - S5:

[0054] S1. Obtain the web page to be extracted, perform feature extraction on the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set.

[0055] In an embodiment of the present invention, the web page to be extracted is a web page determined based on an actual business scenario. For example, in the financial field, the web page to be extracted may be the latest news about various financial products. The web page data feature set refers to the relevant features including the web page content. Among them, the web page data feature set may include category features such as text information features, picture information features, link information features, and non-link information features. The web page data vector set refers to mapping the web page data feature set into a spatial vector to convert the network data into deeper deep semantic information.

[0056] In an embodiment of the present invention, the web page to be extracted can be obtained from a business database (such as a financial database, an insurance database, etc.) by using a preset indexing function (such as Index).

[0057] In an embodiment of the present invention, by performing feature extraction on the web page to be extracted to obtain a web page data feature set, and encoding the web page data feature set to obtain a web page data vector set, the main feature data in the web page to be extracted can be extracted, and some useless words can be removed, which is convenient for improving the efficiency of subsequent web page text extraction.

[0058] As an embodiment of the present invention, the performing feature extraction on the web page to be extracted to obtain a web page data feature set includes:

[0059] Converting the web page to be extracted into a text web page, performing word segmentation processing on the text web page to obtain a word segmentation text set; calculating the weight of each word in the word segmentation text set by using a preset algorithm to obtain the word weight; extracting the words with the word weight greater than a preset threshold from the word segmentation text set as web page keywords; performing part-of-speech tagging on the web page keywords according to a preset dictionary to determine the part-of-speech of the web page keywords; and determining the web page data feature set of the web page to be extracted according to the part-of-speech of the web page keywords.

[0060] Among them, the preset algorithm may be the TFIDF algorithm, and the word weight refers to the frequency of the word appearance. The preset threshold may be 0.75. When the word weight is greater than 0.75, it can be determined that the word is a web page keyword. The part-of-speech tagging refers to finding the corresponding category annotation of the web page keyword in the dictionary. When there is a matching word for the web page keyword in the dictionary, the matching word category annotation is used as the part-of-speech of the web page keyword. Among them, the preset dictionary may be a dictionary customized based on user needs.

[0061] In an embodiment of the present invention, since the parts of speech related to the main content of the web page in the web page keywords are mainly content words such as nouns, verbs, and adjectives, while some function words such as interjections, prepositions, and conjunctions have no actual meaning and contribution to the subsequent determination of the web page type, by determining the part of speech of the web page keywords according to the dictionary, the core feature words in the web page to be extracted can be extracted, and some useless function words can be removed, which can reduce the computational amount of feature extraction and improve the efficiency of feature extraction.

[0062] Further, in the embodiment of the present invention, the Embedding model can be used to perform feature encoding on the web page data feature set to obtain a web page data vector set, which can realize the conversion of the web page data feature set into deeper deep semantic information and further improve the accuracy of web page feature extraction.

[0063] S2. Perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs.

[0064] In the embodiment of the present invention, the indexed web page data set refers to the associated data between the respective data features corresponding to the web page data vector set; the web page type refers to the web page data type, where the web page data type includes text-based web pages, image-based web pages, and link-based web pages.

[0065] In the embodiment of the present invention, by performing a recall process on the web page data vector set to obtain an indexed web page data set, an index can be created for each web page data vector, and by analyzing the classification labels to which the indexed web page data set belongs, the web page type corresponding to the web page to be extracted can be accurately identified.

[0066] As an embodiment of the present invention, the performing a recall process on the web page data vector set to obtain an indexed web page data set includes:

[0067] Obtain the vector labels of the web page data vector set, create partition regions according to the vector labels by using a preset open-source vector database; store the web page data vector set in the partition regions, and create indexes for the web page data vector sets in each partition region to obtain the indexed web page data set.

[0068] Among them, the preset open-source vector database can be Milvus; the partition area can be regarded as a set of vector tags. Corresponding partition areas can be created according to different vector tags. The main function is to avoid storing all web page vector data as a single set. After a large amount of data accumulates in a single set, the query performance will gradually decline. Therefore, by storing the web page data vector set in the partition area, the web page data vector set can be stored in Milvus according to different vector tags, which is convenient for improving the accuracy of indexing subsequently.

[0069] In an embodiment of the present invention, the indexed web page data set can be indexed through the indexing function create_index.

[0070] Further, as Figure 2 shown, determining the web page type corresponding to the web page to be extracted by analyzing the classification tags to which the indexed web page data set belongs includes the following steps S21-S22:

[0071] S21. Obtain the web page data of the web page to be extracted, and select the indexed web page data most similar to the web page data from the indexed web page data set as the pre-classified web page label;

[0072] S22. Select the web page label with the most occurrences in the pre-classified web page labels as the web page type corresponding to the web page to be extracted.

[0073] Among them, the pre-classified web page labels correspond to different web page types. Since web page types include text-based web pages, image-based web pages, link-based web pages, etc., the pre-classified web page labels also include web page labels such as text-based web pages, image-based web pages, and link-based web pages.

[0074] In an embodiment of the present invention, the KNN (K-Nearest Neighbor) algorithm can be used to select the indexed web page data most similar to the web page data from the indexed web page data set as the pre-classified web page label, that is, the process of selecting the most similar indexed web page data is the process of selecting the K value; specifically, different K values correspond to different labels. For example, when K = 3, the corresponding pre-classified web page label is type A (i.e., text-based web page type); when K = 5, the corresponding pre-classified web page label is type B (i.e., image-based web page type); when K = 10, the corresponding pre-classified web page label is type C (i.e., link-based web page type).

[0075] Further, the web page label with the most occurrences in the pre-classified web page labels can be selected as the web page type corresponding to the web page to be extracted. Therefore, if the pre-classified web page label is type C (i.e., link-based web page type), it is the web page type of the web page to be extracted.

[0076] In an alternative embodiment of the present invention, the weight of the indexed web page dataset can also be determined based on the proportion of each data type corresponding to the indexed web page dataset in the indexed web page dataset, where the data type is the same as the web page type and includes text data, image data, link data, etc. For example, if there are ten pieces of data in the indexed web page dataset, including 5 pieces of text data, 3 pieces of image data, and 2 pieces of link data, then the weight corresponding to the text data is 0.5, the weight corresponding to the image data is 0.3, and the weight corresponding to the link data is 0.2. If the proportion of the text data weight exceeds the preset threshold of 0.4, then the web page type corresponding to the indexed web page dataset is a text-based web page.

[0077] S3. Determine whether the web page type of the web page to be extracted is a text-based web page.

[0078] In the implementation of the present invention, the text-based web page refers to a web page dominated by text. For example, in the financial field, news web pages related to bonds and financial products.

[0079] In the embodiment of the present invention, by determining whether the web page type is a text-based web page, it is convenient to apply different methods for extraction to different web page types subsequently.

[0080] In an embodiment of the present invention, the web page label of the web page to be extracted can be queried through a preset query statement (such as an SQL query statement). According to the web page label, the text-based web page type can be determined. Specifically, when the web page label is of type A, it is a text-based web page type; when the web page label is of type B, it is an image-based web page type; when the web page label is of type C, it is a link-based web page type.

[0081] S4. When the web page type of the web page to be extracted is not a text-based web page, extract the web page body of the web page to be extracted according to the preset regular rules.

[0082] In the embodiment of the present invention, the regular rule is a rule for string matching. Among them, the regular rule is also a character pattern composed of ordinary characters and special characters.

[0083] In the embodiment of the present invention, when the web page type is not a text-based web page, it means that the current web page to be extracted belongs to the image-based web page type or the link-based web page type. By extracting the web page body of the web page to be extracted according to the preset regular rules, the web page body extraction can be realized for the image-based web page type or the link-based web page type using the corresponding regular rules. When different web page patterns or web page structures change, there is no need to rewrite the wrapper, which improves the efficiency of web page body extraction.

[0084] As an embodiment of the present invention, the extracting the web page body of the web page to be extracted according to the preset regular rules includes:

[0085] Obtain the web page source code of the web page to be extracted, and determine the position of the web page body in the web page to be extracted according to the web page source code; when it is recognized that the web page to be extracted is an image-type web page, use a preset image regular rule to extract the text from the position of the web page body to obtain the web page body; when it is recognized that the web page to be extracted is a link-type web page, use a preset link regular rule to extract the web page link from the position of the web page body to obtain the web page body.

[0086] Among them, the web page source code refers to the source code of Html; the web page body position refers to the starting position and the ending position of the web page body content, and this web page body position can be obtained by identifying the positions and contents of each tag in Html, generating a DOM tree corresponding to each tag according to the tag, further traversing the DOM tree to obtain the node path information corresponding to each tag, and obtaining the web page body position corresponding to Html according to the node path information.

[0087] Specifically, the positions of each Html tag can be , and etc. Among them, different tags may include start tags and end tags at different positions. If there is content between the start tag and the end tag, the content is assigned to the tag; each tag is regarded as a node, and according to the distribution logic of the web page, each tag is connected to form a DOM tree; by traversing the DOM tree, the node path information of each tag is obtained, such as: ,, 、 、 , etc. According to this path information, the starting position and the ending position of the corresponding web page body can be determined.

[0088] In an embodiment of the present invention, the strings in the image-type web page can be matched at the beginning and the end through an image regular rule to extract the web page body. Specifically, the image regular rule can be allfinds = ( \(+?)), where allfinds can be the web page title, can be the starting position of the web page body, can be the starting position of the image in the body, +? is a lazy qualifier, indicating that the qualifier in the regular rule can be repeated 1 time or more times, can be the ending position of the web page body, and can be the ending position of the image in the body.

[0089] Furthermore, the link regular rule is similar to the image regular rule, and only the character limit of the regular rule needs to be modified according to the requirements, which will not be elaborated here.

[0090] S5. When the web page type of the to-be-extracted web page is a text-type web page, use the trained web page body extraction model to extract the web page body of the to-be-extracted web page.

[0091] In an embodiment of the present invention, the trained web page body extraction model can be constructed by a BiLSTM model, where BiLSTM is a bidirectional long short-term memory network.

[0092] In an alternative embodiment of the present invention, after the web page type of the to-be-extracted web page is a text-type web page, a DOM tree of the web page can be constructed, and the DOM tree can be traversed through Babel to determine the paths of each tag node, obtain the web page body position, and remove invalid node information such as web page links and scripts in the web page body. By using blank lines to replace invalid nodes such as web page links and scripts, the remaining text information of the web page is retained, thereby forming a sequence of line data. Finally, the processed data is input into the trained web page body extraction model, and the model is used to identify which lines in the distribution of text lines and blank lines in the web page body belong to the real text lines of the web page body, and the text information is extracted.

[0093] In an embodiment of the present invention, when the web page type is a text-type web page, use the trained web page body extraction model to extract the web page body of the to-be-extracted web page, which can directly and accurately extract the web page body of the to-be-extracted web page through the model, avoid extracting some content unrelated to the web page body during the extraction process, and improve the accuracy of web page body extraction.

[0094] As an embodiment of the present invention, as Figure 3 shown, the extracting the web page body of the to-be-extracted web page by using the trained web page body extraction model includes the following steps S51-S53:

[0095] S51. Encode the web page to be extracted using the bidirectional long short-term memory network in the trained web page body extraction model to obtain an encoded data set;

[0096] S52. Decode the encoded data set using the unidirectional long short-term memory network in the web page body extraction model to obtain a decoded data set;

[0097] S53. Input the decoded data set into a preset activation function to obtain an activation probability value, and obtain the web page body according to the activation probability value.

[0098] Among them, by encoding using the bidirectional long short-term memory network in the trained web page body extraction model, then decoding the encoded data using the unidirectional long short-term memory network, and obtaining the probability of falling on each interval through the activation function after dimensional compression, the web page body is obtained according to the activation probability. Among them, the activation function is the softmax function.

[0099] Further, the encoding of the web page to be extracted using the bidirectional long short-term memory network in the trained web page body extraction model to obtain an encoded data set includes:

[0100] Calculate the state value of the web page to be extracted using the input gate in the bidirectional long short-term memory network; calculate the activation value of the web page to be extracted using the forget gate in the bidirectional long short-term memory network; calculate the state update value of the web page to be extracted according to the state value and the activation value; calculate the encoded data set corresponding to the state update value using the output gate in the bidirectional long short-term memory network.

[0101] In an alternative embodiment of the present invention, the calculation method of the state value includes:

[0102]

[0103] where i t represents the state value, represents the bias of the cell unit in the input gate, w i represents the activation factor of the input gate, h t-1 represents the peak value of the web page to be extracted at the t-1 moment of the input gate, x t represents the web page to be extracted at the t moment, b i represents the weight of the cell unit in the input gate.

[0104] In an alternative embodiment of the present invention, the calculation method of the activation value includes:

[0105]

[0106] Among them, f t represents the activation value, represents the bias of the cell unit in the forgetting gate, w f represents the activation factor of the forgetting gate, represents the peak value of the web page to be extracted at the (t - 1)th moment in the forgetting gate, x t represents the web page to be extracted input at the tth moment, b f represents the weight of the cell unit in the forgetting gate.

[0107] In an alternative embodiment of the present invention, the calculation method of the state update value includes:

[0108]

[0109] Among them, c t represents the state update value, h t-1 represents the peak value of the web page to be extracted at the (t - 1)th moment in the input gate, represents the peak value of the web page to be extracted at the (t - 1)th moment in the forgetting gate.

[0110] In an alternative embodiment of the present invention, calculating the encoded data set corresponding to the state update value by using the output gate in the bidirectional long short - term memory network includes:

[0111] Calculating the encoded data set by using the following formula:

[0112] o t = tanh(c t )

[0113] Among them, o t represents the encoded data set, tanh represents the activation function of the output gate, c t represents the state update value.

[0114] In the embodiment of the present invention, first, by extracting the features of the web page to be extracted, a web page data feature set is obtained, and the web page data feature set is encoded to obtain a web page data vector set, so that the main feature data in the web page to be extracted can be extracted and some useless words can be removed, facilitating the improvement of the efficiency of subsequent web page text extraction; secondly, by performing a recall process on the web page data vector set, an indexed web page data set is obtained, an index can be created for each web page data vector, and by analyzing the classification labels to which the indexed web page data set belongs, the web page type corresponding to the web page to be extracted can be accurately identified, facilitating the subsequent application of different methods for extraction for different web page types; finally, when the identified web page type is not a text-based web page, the web page text is extracted by using regular rules, and when the identified web page type is a text-based web page, the trained web page text extraction model is used to extract the web page text of the web page to be extracted, which can avoid extracting some content unrelated to the web page text during the extraction process, improve the accuracy of web page text extraction, and different web page types can apply different methods for targeted extraction of web page text. When different web page modes or web page structures change, there is no need to rewrite the wrapper, improving the efficiency of web page text extraction. Therefore, the web page text extraction method proposed in the embodiment of the present invention can improve the efficiency and accuracy of web page text extraction.

[0115] The web page text extraction device 100 according to the present invention can be installed in an electronic device. According to the functions achieved, the web page text extraction device may include a web page feature extraction module 101, a web page type recognition module 102, and a web page text extraction module 103. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0116] In this embodiment, the functions of each module / unit are as follows:

[0117] The web page feature extraction module 101 is used to obtain the web page to be extracted, extract the features of the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set.

[0118] In the embodiment of the present invention, the web page to be extracted is a web page determined based on an actual business scenario. For example, in the financial field, the web page to be extracted can be the latest news about various financial products; the web page data feature set refers to the relevant features including the web page content, where the web page data feature set can include category features such as text information features, picture information features, link information features, and non-link information features; the web page data vector set refers to mapping the web page data feature set into a space vector to convert network data into deeper deep semantic information.

[0119] In one embodiment of the present invention, the web page to be extracted can obtain the web page to be extracted from a service database (such as a financial database, an insurance database, etc.) by using a preset indexing function (such as Index).

[0120] In an embodiment of the present invention, by performing feature extraction on the web page to be extracted to obtain a web page data feature set, and encoding the web page data feature set to obtain a web page data vector set, the main feature data in the web page to be extracted can be extracted, and some useless words can be removed, which is convenient for improving the efficiency of subsequent web page text extraction.

[0121] As an embodiment of the present invention, the web page feature extraction module 101 performs the following operations to perform feature extraction on the web page to be extracted to obtain a web page data feature set, including:

[0122] Convert the web page to be extracted into a text web page, and perform word segmentation processing on the text web page to obtain a word segmentation text set;

[0123] Use a preset algorithm to calculate the weight of each word in the word segmentation text set to obtain a word weight;

[0124] Extract the words in the word segmentation text set whose word weights are greater than a preset threshold as web page keywords;

[0125] Perform part-of-speech tagging on the web page keywords according to a preset dictionary to determine the part-of-speech of the web page keywords;

[0126] Determine the web page data feature set of the web page to be extracted according to the part-of-speech of the web page keywords.

[0127] Among them, the preset algorithm can be the TFIDF algorithm, and the word weight refers to the frequency of the word; the preset threshold can be 0.75, and when the word weight is greater than 0.75, it can be determined that the word is a web page keyword; the part-of-speech tagging refers to finding the corresponding category annotation of the web page keyword in the dictionary, and when there is a matching word for the web page keyword in the dictionary, the matching word category annotation is used as the part-of-speech of the web page keyword. Among them, the preset dictionary can be a dictionary customized based on user needs.

[0128] In one embodiment of the present invention, since the parts of speech related to the main content of the web page in the web page keywords are mainly content words such as nouns, verbs, and adjectives, while some function words such as interjections, prepositions, and conjunctions have no actual meaning and contribution to subsequent determination of the web page type, so by performing part-of-speech tagging according to the dictionary to determine the part-of-speech of the web page keywords, the core feature words in the web page to be extracted can be extracted, and some useless function words can be removed, which can reduce the calculation amount of feature extraction and improve the efficiency of feature extraction.

[0129] Furthermore, in the embodiments of the present invention, an Embedding model can be used to perform feature encoding on the web page data feature set to obtain a web page data vector set, which can realize the conversion of the web page data feature set into deeper deep semantic information and further improve the accuracy of web page feature extraction.

[0130] The web page type recognition module 102 is configured to perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs.

[0131] In the embodiments of the present invention, the indexed web page data set refers to the associated data between the respective data features corresponding to the web page data vector set; the web page type refers to the web page data type, where the web page data type includes text-based web pages, image-based web pages, and link-based web pages.

[0132] In the embodiments of the present invention, by performing a recall process on the web page data vector set to obtain an indexed web page data set, an index can be created for each web page data vector, and by analyzing the classification labels to which the indexed web page data set belongs, the web page type corresponding to the web page to be extracted can be accurately identified.

[0133] As an embodiment of the present invention, the web page type recognition module 102 performs a recall process on the web page data vector set to obtain an indexed web page data set by performing the following operations, including:

[0134] Obtain the vector labels of the web page data vector set, and create partition regions according to the vector labels by using a preset open-source vector database;

[0135] Store the web page data vector set into the partition regions, and create indexes for the web page data vector sets in each of the partition regions to obtain the indexed web page data set.

[0136] Wherein, the preset open-source vector database can be Milvus; the partition regions can be regarded as a set of vector labels, and corresponding partition regions can be created according to different vector labels. The main function is to avoid storing all web page vector data as a single set. When a set accumulates a large amount of data, the query performance will gradually decline. Therefore, by storing the web page data vector set into the partition regions, the web page data vector set can be stored in Milvus according to different vector labels, which is convenient for improving the accuracy of indexes subsequently.

[0137] In an embodiment of the present invention, the indexed web page data set can be indexed by using the index function create_index.

[0138] Further, determining the web page type corresponding to the web page to be extracted by analyzing the classification tags to which the indexed web page dataset belongs includes:

[0139] Obtain the web page data of the web page to be extracted, and select the indexed web page data most similar to the web page data from the indexed web page dataset as the pre-classified web page tag; select the web page tag with the most occurrences in the pre-classified web page tags as the web page type corresponding to the web page to be extracted.

[0140] Among them, the pre-classified web page tags correspond to different web page types. Since the web page types include text-based web pages, image-based web pages, link-based web pages, etc., the pre-classified web page tags also include web page tags such as text-based web pages, image-based web pages, and link-based web pages.

[0141] In an embodiment of the present invention, the KNN (K-Nearest Neighbor) algorithm can be used to select the indexed web page data most similar to the web page data from the indexed web page dataset as the pre-classified web page tag, that is, the process of selecting the most similar indexed web page data and the process of selecting the K value; specifically, different K values correspond to different tags. For example, when K = 3, the corresponding pre-classified web page tag is type A (i.e., text-based web page type); when K = 5, the corresponding pre-classified web page tag is type B (i.e., image-based web page type); when K = 10, the corresponding pre-classified web page tag is type C (i.e., link-based web page type).

[0142] Further, the web page tag with the most occurrences in the pre-classified web page tags can be selected as the web page type corresponding to the web page to be extracted. Therefore, if the pre-classified web page tag is type C (i.e., link-based web page type), it is the web page type of the web page to be extracted.

[0143] In an alternative embodiment of the present invention, the weight of the indexed web page dataset can also be determined based on the proportion of each data type corresponding to the indexed web page dataset in the indexed web page dataset. Among them, the data type is the same as the web page type, including text data, image data, link data, etc. For example, there are ten pieces of data in the indexed web page dataset, including 5 pieces of text data, 3 pieces of image data, and 2 pieces of link data. Then the weight corresponding to the text data is 0.5, the weight corresponding to the image data is 0.3, and the weight corresponding to the link data is 0.2. If the weight proportion of the text data exceeds the preset threshold of 0.4, the web page type corresponding to the indexed web page dataset is a text-based web page.

[0144] The web page body extraction module 103 is used to determine whether the web page type of the web page to be extracted is a text-based web page; when the web page type of the web page to be extracted is not a text-based web page, the web page body of the web page to be extracted is extracted according to a preset regular rule; when the web page type of the web page to be extracted is a text-based web page, the web page body of the web page to be extracted is extracted by using the trained web page body extraction model.

[0145] In the implementation of the present invention, the text-based web page refers to a web page dominated by text. For example, in the financial field, news web pages related to bonds and financial products.

[0146] In the embodiment of the present invention, by determining whether the web page type is a text-based web page, it is convenient to apply different methods for extraction to different web page types subsequently.

[0147] In an embodiment of the present invention, the web page tags of the web page to be extracted can be queried through a preset query statement (such as an SQL query statement). According to the web page tags, the text-based web page type can be determined. Specifically, when the web page tag is of type A, it is a text-based web page type; when the web page tag is of type B, it is an image-based web page type; when the web page tag is of type C, it is a link-based web page type.

[0148] In the embodiment of the present invention, the regular rule is a rule for string matching. Among them, the regular rule is also a character pattern composed of ordinary characters and special characters.

[0149] In the embodiment of the present invention, when the web page type is not a text-based web page, it means that the current web page to be extracted belongs to the image-based web page type or the link-based web page type. By extracting the web page body of the web page to be extracted according to the preset regular rule, the corresponding regular rule can be used to implement web page body extraction for the image-based web page type or the link-based web page type. When different web page modes or web page structures change, there is no need to rewrite the wrapper, which improves the efficiency of web page body extraction.

[0150] As an embodiment of the present invention, the web page body extraction module 103 extracts the web page body of the web page to be extracted according to the preset regular rule by performing the following operations, including:

[0151] Obtain the web page source code of the web page to be extracted, and determine the position of the web page body in the web page according to the web page source code;

[0152] When it is recognized that the web page to be extracted is an image-based web page, use the preset image regular rule to extract the text from the web page body position to obtain the web page body;

[0153] When identifying that the web page to be extracted is a link-type web page, use a preset link regular rule to extract web page links from the web page body position to obtain the web page body.

[0154] Among them, the web page source code refers to the source code of Html; the web page body position refers to the starting position and the ending position of the web page body content, and this web page body position can be obtained by identifying the positions and contents of each tag in Html, generating a DOM tree corresponding to each tag according to the tag, further traversing the DOM tree to obtain the node path information corresponding to each tag, and obtaining the web page body position corresponding to Html according to the node path information.

[0155] Specifically, the positions of each tag in Html can be , and etc. Among them, different tags may include start tags and end tags at different positions. If there is content between the start tag and the end tag, the content is assigned to the tag; each tag is regarded as a node, and according to the distribution logic of the web page, each tag is connected to form a DOM tree; by traversing the DOM tree, the node path information of each tag is obtained, such as:,, 、 、 , etc. According to this path information, the starting position and ending position of the corresponding web page body can be determined.

[0156] In an embodiment of the present invention, the strings in the image-type web page can be matched at the beginning and end through an image regular rule to extract the web page body. Specifically, the image regular rule can be allfinds = ( \(+?)), where allfinds can be the web page title, can be the starting position of the web page body, can be the starting position of the image in the body, +? is a lazy qualifier, indicating that the qualifier in the regular rule can be repeated 1 time or more times, can be the ending position of the web page body, and can be the ending position of the image in the body.

[0157] Furthermore, the link regular rule is similar to the image regular rule, and only the character limit of the regular rule needs to be modified according to the requirements, which will not be elaborated here.

[0158] In an embodiment of the present invention, the trained web page body extraction model can be constructed by a BiLSTM model, where BiLSTM is a bidirectional long short-term memory network.

[0159] In an alternative embodiment of the present invention, after the web page type of the web page to be extracted is a text-type web page, a DOM tree of the web page can be constructed, and the DOM tree can be traversed through Babel to determine the paths of each tag node, obtain the web page body position, and remove invalid node information such as web page links and scripts in the web page body. By using blank lines to replace invalid nodes such as web page links and scripts, the remaining text information of the web page is retained, thereby forming a sequence of line data. Finally, the processed data is input into the trained web page body extraction model, and the model is used to identify which lines in the distribution of text lines and blank lines in the web page body belong to the real text lines of the body, and the body information is extracted.

[0160] In the embodiment of the present invention, when the web page type is a text-type web page, the trained web page body extraction model is used to extract the web page body of the web page to be extracted, and the web page body of the web page to be extracted can be accurately extracted directly through the model, avoiding extracting some content unrelated to the web page body during the extraction process, and improving the accuracy of web page body extraction.

[0161] As an embodiment of the present invention, the web page body extraction module 103 can also be used to perform the following operations to extract the web page body of the web page to be extracted by using the trained web page body extraction model, including:

[0162] Encoding the web page to be extracted by using the bidirectional long short-term memory network in the trained web page body extraction model to obtain an encoded data set;

[0163] The encoded data set is decoded using the single-direction long short-term memory network in the web page body extraction model to obtain a decoded data set;

[0164] The decoded data set is input into a preset activation function to obtain an activation probability value, and the web page body is obtained according to the activation probability value.

[0165] Among them, encoding is performed using the bidirectional long short-term memory network in the trained web page body extraction model, and then the encoded data is decoded using the single-direction long short-term memory network, and the probability of falling on each interval is obtained through the activation function after dimensional compression. The web page body is obtained according to the activation probability. Among them, the activation function is the softmax function.

[0166] Further, encoding the web page to be extracted using the bidirectional long short-term memory network in the trained web page body extraction model to obtain an encoded data set includes:

[0167] Calculating the state value of the web page to be extracted using the input gate in the bidirectional long short-term memory network; calculating the activation value of the web page to be extracted using the forget gate in the bidirectional long short-term memory network; calculating the state update value of the web page to be extracted according to the state value and the activation value; calculating the encoded data set corresponding to the state update value using the output gate in the bidirectional long short-term memory network.

[0168] In an alternative embodiment of the present invention, the calculation method of the state value includes:

[0169]

[0170] where i t represents the state value, represents the bias of the cell unit in the input gate, w i represents the activation factor of the input gate, h t-1 represents the peak value of the web page to be extracted at the t-1 moment of the input gate, x t represents the web page to be extracted at the t moment, b i represents the weight of the cell unit in the input gate.

[0171] In an alternative embodiment of the present invention, the calculation method of the activation value includes:

[0172]

[0173] where f t represents the activation value, represents the bias of the cell unit in the forget gate, w f represents the activation factor of the forget gate, represents the peak value of the web page to be extracted at the forgetting gate t-1 moment, x t represents the web page to be extracted input at the t moment, b f represents the weight of the cell unit in the forgetting gate.

[0174] In an alternative embodiment of the present invention, the calculation method of the state update value includes:

[0175]

[0176] where c t represents the state update value, h t-1 represents the peak value of the web page to be extracted at the input gate t-1 moment, represents the peak value of the web page to be extracted at the forgetting gate t-1 moment.

[0177] In an alternative embodiment of the present invention, calculating the encoded data set corresponding to the state update value by using the output gate in the bidirectional long short-term memory network includes:

[0178] Calculating the encoded data set by using the following formula:

[0179] o t = tanh(c t )

[0180] where o t represents the encoded data set, tanh represents the activation function of the output gate, c t represents the state update value.

[0181] In the embodiment of the present invention, first, by performing feature extraction on the to-be-extracted web page, a web page data feature set is obtained, and the web page data feature set is encoded to obtain a web page data vector set, so that the main feature data in the to-be-extracted web page can be extracted and some useless words can be removed, facilitating the improvement of the efficiency of subsequent web page body extraction; secondly, by performing a recall process on the web page data vector set, an indexed web page data set is obtained, an index can be created for each web page data vector, and by analyzing the classification labels to which the indexed web page data set belongs, the web page type corresponding to the to-be-extracted web page can be accurately identified, facilitating the subsequent application of different methods for extraction for different web page types; finally, when the identified web page type is not a text-based web page, the web page body is extracted by using regular rules, and when the identified web page type is a text-based web page, the trained web page body extraction model is used to extract the web page body of the to-be-extracted web page, which can avoid extracting some content irrelevant to the web page body during the extraction process, improve the accuracy of web page body extraction, and different web page types can apply different methods for targeted extraction of the web page body. When different web page modes or web page structures change, there is no need to rewrite the wrapper, improving the efficiency of web page body extraction. Therefore, the web page body extraction device based on the embodiment of the present invention can improve the efficiency and accuracy of web page body extraction.

[0182] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the web page body extraction method of the present invention.

[0183] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a web page body extraction program.

[0184] Among them, the memory 11 includes at least one type of medium, which includes flash memory, external hard drive, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, local disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device, such as the external hard drive of the electronic device. In some other embodiments, the memory 11 can also be an external storage device of the electronic device, such as a plug-in external hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device. Further, the memory 11 can also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of the web page text extraction program, etc., but also to temporarily store data that has been output or will be output.

[0185] In some embodiments, the processor 10 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules (such as the web page text extraction program, etc.) stored in the memory 11, and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.

[0186] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The communication bus 12 is set to enable connection communication between the memory 11 and at least one processor 10, etc. For the convenience of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0187] Figure 5 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 5 The structures shown do not constitute a limitation on the electronic device, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0188] For example, although not shown, the electronic device may further include a power source (such as a battery) for powering each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0189] Optionally, the communication interface 13 may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices.

[0190] Optionally, the communication interface 13 may further include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0191] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0192] The web page text extraction program stored in the memory 11 of the electronic device is a combination of multiple computer programs. When running in the processor 10, it can implement:

[0193] Obtain the web page to be extracted, extract the features of the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set;

[0194] Perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs;

[0195] Determine whether the web page type of the web page to be extracted is a text-based web page;

[0196] When the web page type of the web page to be extracted is not a text-based web page, extract the web page body of the web page to be extracted according to a preset regular rule;

[0197] When the web page type of the web page to be extracted is a text-based web page, use the trained web page body extraction model to extract the web page body of the web page to be extracted.

[0198] Specifically, for the specific implementation method of the above computer program by the processor 10, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0199] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable medium. The computer-readable medium can be non-volatile or volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0200] An embodiment of the present invention can also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor of an electronic device, it can implement:

[0201] Obtain a web page to be extracted, extract features of the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set;

[0202] Perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs;

[0203] Determine whether the web page type of the web page to be extracted is a text-based web page;

[0204] When the web page type of the web page to be extracted is not a text-based web page, extract the web page body of the web page to be extracted according to a preset regular rule;

[0205] When the web page type of the web page to be extracted is a text-based web page, use the trained web page body extraction model to extract the web page body of the web page to be extracted.

[0206] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0207] In several embodiments provided by the present invention, it should be understood that the disclosed medium, device, apparatus, and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0208] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0209] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0210] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0211] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.

[0212] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0213] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular form does not exclude the plural form. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Terms such as "second" are used to denote names and do not denote any particular order.

[0214] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for extracting web page text, characterized in that, the method includes: Obtain the web page to be extracted, perform feature extraction on the web page to be extracted to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set; Perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs; Judge whether the web page type of the web page to be extracted is a text-based web page; When the web page type of the web page to be extracted is not a text-based web page, extract the web page text of the web page to be extracted according to a preset regular rule; When the web page type of the web page to be extracted is a text-based web page, use the trained web page text extraction model to extract the web page text of the web page to be extracted; Among them, the performing a recall process on the web page data vector set to obtain an indexed web page data set includes: obtaining the vector labels of the web page data vector set, creating partition regions according to the vector labels by using a preset open-source vector database; storing the web page data vector set in the partition regions, and creating indexes for the web page data vector sets in each partition region to obtain the indexed web page data set; The performing feature extraction on the web page to be extracted to obtain a web page data feature set includes: converting the web page to be extracted into a text web page, performing word segmentation processing on the text web page to obtain a word segmentation text set; calculating the weight of each word in the word segmentation text set by using a preset algorithm to obtain word weights; extracting the words with word weights greater than a preset threshold from the word segmentation text set as web page keywords; performing part-of-speech tagging on the web page keywords according to a preset dictionary to determine the part-of-speech of the web page keywords; determining the web page data feature set of the web page to be extracted according to the part-of-speech of the web page keywords.

2. The method for extracting web page text according to claim 1, characterized in that, the determining the web page type corresponding to the web page to be extracted by analyzing the classification labels to which the indexed web page data set belongs includes: Obtaining the web page data of the web page to be extracted, and selecting the indexed web page data most similar to the web page data from the indexed web page data set as a pre-classified web page label; Selecting the web page label with the most occurrences in the pre-classified web page labels as the web page type corresponding to the web page to be extracted.

3. The method for extracting web page text according to claim 1, characterized in that, the using the trained web page text extraction model to extract the web page text of the web page to be extracted includes: Encoding the web page to be extracted by using a bidirectional long short-term memory network in the trained web page text extraction model to obtain an encoded data set; Performing decoding processing on the encoded data set by using a unidirectional long short-term memory network in the web page text extraction model to obtain a decoded data set; Inputting the decoded data set into a preset activation function to obtain an activation probability value, and obtaining the web page text according to the activation probability value.

4. The method for extracting web page text according to claim 3, It is characterized in that encoding the to-be-extracted web page by using a bidirectional long short-term memory network in the trained web page body extraction model to obtain an encoded data set, including: calculating a state value of the to-be-extracted web page by using an input gate in the bidirectional long short-term memory network; calculating an activation value of the to-be-extracted web page by using a forgetting gate in the bidirectional long short-term memory network; calculating a state update value of the to-be-extracted web page according to the state value and the activation value; calculating an encoded data set corresponding to the state update value by using an output gate in the bidirectional long short-term memory network.

5. The web page body extraction method according to claim 1, It is characterized in that extracting the web page body of the to-be-extracted web page according to a preset regular rule includes: obtaining the web page source code of the to-be-extracted web page, and determining the web page body position in the to-be-extracted web page according to the web page source code; when identifying that the to-be-extracted web page is an image-type web page, extracting the web page body from the web page body position by using a preset image regular rule to obtain the web page body; when identifying that the to-be-extracted web page is a link-type web page, extracting the web page link from the web page body position by using a preset link regular rule to obtain the web page body.

6. A web page body extraction device for implementing the web page body extraction method according to any one of claims 1 to 5, It is characterized in that the device includes: a web page feature extraction module, configured to obtain a to-be-extracted web page, extract features of the to-be-extracted web page to obtain a web page data feature set, and encode the web page data feature set to obtain a web page data vector set; a web page type identification module, configured to perform a recall process on the web page data vector set to obtain an indexed web page data set, and determine the web page type corresponding to the to-be-extracted web page by analyzing classification tags to which the indexed web page data set belongs; a web page body extraction module, configured to determine whether the web page type of the to-be-extracted web page is a text-type web page; when the web page type of the to-be-extracted web page is not a text-type web page, extracting the web page body of the to-be-extracted web page according to a preset regular rule; when the web page type of the to-be-extracted web page is a text-type web page, extracting the web page body of the to-be-extracted web page by using a trained web page body extraction model.

7. An electronic device, It is characterized in that the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the web page body extraction method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, It is characterized in that when the computer program is executed by a processor, it implements the web page body extraction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Website information query method and system thereof

    CN101673306A

  • Method for inputting and processing feature word in file content

    WO2011006412A1