A method, device and storage medium for extracting the main text of a tender web page
By constructing a DOM tree and combining convolutional neural network model and rule screening, the problem of inefficient extraction of text on the bidding web page is solved, and efficient and accurate text extraction is achieved.
Patent Information
- Application Number
- CN202210163751.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-02-22
AI Technical Summary
The existing methods of tender web page text extraction are inefficient and prone to error analysis. The traditional features are not suitable for complex and changeable tender web pages. Machine learning-based methods take a long time and are difficult to judge the boundaries of the web page text.
The bidding web page body is extracted into the optimal path search problem, a DOM tree is constructed, and the node with the highest node score and the longest text length is determined. Combined with the convolutional neural network model and rule filtering, the target text is obtained.
It improves the efficiency and accuracy of the text extraction of bidding web pages, narrows the search space, integrates traditional features, deep learning algorithms and rule screening methods, and improves the accuracy of text extraction.
Smart Images

Figure CN115098812B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing. More specifically, the present application relates to a method, apparatus, and storage medium for extracting the main text of a tender web page. Background Art
[0002] In recent years, more and more websites have started to publish various tender information. By obtaining effective tender information in a timely manner, enterprises can gain many benefits. However, the tender information published on each website uses different layouts, which are filled with a large number of invalid information such as navigation bars and advertisements. These invalid information will prevent users from focusing on the key information immediately, resulting in a waste of a large amount of time in the review of a large number of tender texts. At the same time, these invalid information will also make the recognition of downstream tasks more difficult.
[0003] Traditional methods for extracting the main text of a tender web page generally parse each tender site by professional researchers. This method requires a huge amount of human resources, has low efficiency, and may also result in incorrect parsing, causing all the tender main texts of a website to be parsed incorrectly. In addition, there are also methods for extracting the main text of a web page based on statistical web page features such as text block density and tag path features, but these features are not applicable to tender web pages. Summary of the Invention
[0004] Based on the above technical problems, the present invention aims to transform the problem of extracting the main text of a tender web page into an optimal path search problem, fuse traditional features and deep learning algorithms for efficient search, and use a rule screening method to obtain the target main text, so as to improve the accuracy of extracting the main text of a tender web page.
[0005] The first aspect of the present invention provides a method for extracting the main text of a tender web page, the method comprising:
[0006] Construct a DOM tree for the tender web page to be extracted;
[0007] Determine a first node with the highest node score and a second node with the longest text length in the current level of the DOM tree;
[0008] Determine the optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node in a text set to be screened, where the text set to be screened includes texts of optimal nodes corresponding to multiple levels;
[0009] Perform rule screening on the text set to be screened to obtain the target main text.
[0010] In some embodiments of the present invention, the determining the optimal node corresponding to the current level from the first node and the second node includes:
[0011] If the first node is the same as the second node, determine the first node as the optimal node;
[0012] If the first node is not the same as the second node, select the optimal node from the first node and the second node based on a preset convolutional neural network model.
[0013] In some embodiments of the present invention, after determining the first node as the optimal node if the first node is the same as the second node, it further includes:
[0014] Determine the level formed by the child nodes of the first node as the current level, and repeatedly execute the step of determining the first node with the highest node score and the second node with the longest text length in the currently determined DOM tree level to determine the optimal nodes corresponding to the multiple levels.
[0015] In some embodiments of the present invention, selecting the optimal node from the first node and the second node based on a preset convolutional neural network model includes:
[0016] Input the texts corresponding to the first node and the second node into the preset convolutional neural network model, and output the text classification results corresponding to the first node and the second node;
[0017] Select the optimal node from the first node and the second node according to the text classification results corresponding to the first node and the second node.
[0018] In some embodiments of the present invention, the selecting the optimal node from the first node and the second node according to the text classification results corresponding to the first node and the second node includes:
[0019] If the text classification results corresponding to both the first node and the second node are non-main texts, select the node with the largest number of p tags as the optimal node;
[0020] If the text classification results corresponding to both the first node and the second node are not non-main texts, select the node with the smallest probability among the non-main text tags as the optimal node;
[0021] If the text classification result corresponding to the first node is non-main text, select the second node as the optimal node;
[0022] If the text classification result corresponding to the second node is non-main text, select the first node as the optimal node.
[0023] In some embodiments of the present invention, determining the first node with the highest node score in the currently determined DOM tree level includes:
[0024] Calculate the node scores of all nodes at the current level of the DOM tree based on the text density and symbol density;
[0025] Select the node with the highest node score from all nodes at the current level as the first node.
[0026] In some embodiments of the present invention, the calculating the node scores of all nodes at the current level of the DOM tree based on the text density and symbol density, the formula is:
[0027]
[0028] Wherein, td represents the text density of the node, sbd represents the symbol density of the node, p represents the number of p tags, ntd represents the set of text densities at the current level, np represents the set of the number of p tags at the current level, and nsbd represents the set of symbol densities at the current level.
[0029] In some embodiments of the present invention, the performing rule screening on the text set to be screened to obtain the target body text includes:
[0030] Denote the label predicted by the convolutional neural network model as the body text label as the Y label; take the longest text corresponding to the Y label as the initial solution;
[0031] Perform rule screening on the initial solution based on the text length ratio and the link text ratio to obtain the target body text.
[0032] The second aspect of the present invention provides a device for extracting the body text of a tender web page, the device includes:
[0033] A construction module, configured to construct a DOM tree for the tender web page to be extracted;
[0034] A determination module, configured to determine a first node with the highest node score and a second node with the longest text length in the current level of the DOM tree;
[0035] A comparison module, configured to determine the optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node into the text set to be screened, where the text set to be screened includes the texts of the optimal nodes corresponding to multiple levels;
[0036] A screening module, configured to perform rule screening on the text set to be screened to obtain the target body text.
[0037] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0038] Construct a DOM tree for the tender web page to be extracted;
[0039] Determine the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree;
[0040] Determine the optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node in the text set to be screened. The text set to be screened includes the texts of the optimal nodes corresponding to multiple levels;
[0041] Perform rule screening on the text set to be screened to obtain the target main text.
[0042] A fourth aspect of the present invention provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0043] Construct a DOM tree for the tender web page to be extracted;
[0044] Determine the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree;
[0045] Determine the optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node in the text set to be screened. The text set to be screened includes the texts of the optimal nodes corresponding to multiple levels;
[0046] Perform rule screening on the text set to be screened to obtain the target main text.
[0047] The beneficial effects of the present application are as follows: The present application constructs a DOM tree for the tender web page to be extracted, determines the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree, determines the optimal node corresponding to the current level from the first node and the second node, converts the method for extracting the main text of the tender web page into an optimal path search problem, greatly improves the efficiency and reduces the space; and stores the text corresponding to the optimal node in the text set to be screened. The text set to be screened includes the texts of the optimal nodes corresponding to multiple levels, and performs rule screening on the text set to be screened to obtain the target main text. Here, the traditional features, deep learning algorithms and rule screening methods are integrated to obtain the target main text, thereby improving the accuracy of extracting the main text.
[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings forming a part of the specification depict the embodiments of the present application and are used together with the description to explain the principles of the present application.
[0050] With reference to the accompanying drawings, the present application can be more clearly understood according to the following detailed description, where:
[0051] Figure 1 Fig. shows a schematic diagram of the steps of a method for extracting the text body of a tender web page according to an exemplary embodiment of the present application;
[0052] Figure 2 Fig. shows a flowchart of the process of selecting the optimal node according to an exemplary embodiment of the present application;
[0053] Figure 3 Fig. shows a flowchart of rule screening according to an exemplary embodiment of the present application;
[0054] Figure 4 Fig. shows a schematic structural diagram of an apparatus for extracting the text body of a tender web page according to an exemplary embodiment of the present application;
[0055] Figure 5 Fig. shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application;
[0056] Figure 6 Fig. shows a schematic diagram of a storage medium provided by an exemplary embodiment of the present application. Detailed Embodiments
[0057] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application. It is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other examples, some technical features well known to those skilled in the art are not described to avoid confusion with the present application.
[0058] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or combinations thereof.
[0059] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. The drawings are not drawn to scale, and some details may be enlarged for the purpose of clear expression, and some details may be omitted. The shapes of various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0060] The following will describe several embodiments in conjunction with the appended Figure 1-6 drawings to describe the exemplary embodiments according to the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principles of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0061] Traditional methods for extracting the main text of tender web pages involve professional researchers parsing each tender website. This approach consumes a huge amount of human resources, is inefficient, and may also result in incorrect parsing, causing all the main text of a website to be misparsed. In other methods, methods for extracting the main text of web pages based on statistical web page features such as text block density and tag path features are only applicable to web pages with regular structural features such as news, and are not applicable to complex and variable tender web pages. Moreover, this method is prone to misjudging invalid information such as top / bottom navigation bars, floating windows, QR codes, and page-turning links as the main text. In other methods, there are also methods for extracting the main text of web pages based on machine learning, which judge the main text through text classification. This method requires judging the text represented by all nodes in the web page DOM tree, consuming a long time; at the same time, since there are multiple confusing child nodes in the main text nodes, it is difficult for the above method to determine the boundary of the web page main text, thus misjudging the child nodes as main text nodes.
[0062] Therefore, in some exemplary embodiments of the present application, a method for extracting the main text of a tender web page is provided, as Figure 1 shown, the method includes:
[0063] S1. Construct a DOM tree for the tender web page to be extracted;
[0064] S2. Determine the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree;
[0065] S3. Determine the optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node into the text set to be screened. The text set to be screened includes texts of optimal nodes corresponding to multiple levels;
[0066] S4. Perform rule screening on the text set to be screened to obtain the target main text.
[0067] In a specific implementation, after constructing the DOM tree for the tender web page to be extracted, it is also necessary to preprocess the DOM tree to reduce the number of nodes to be compared, thereby increasing the search efficiency and avoiding the confusion caused by some "dirty" nodes. Specifically, 3 preprocessing means can be adopted: removing tags that have no text and no sub-tags; removing IDs that definitely do not contain the main text ("footer", "copyright"); removing classes that definitely do not contain the main text ("footer", "copyright", "foot").
[0068] In a preferred implementation manner, determining the optimal node corresponding to the current level from the first node and the second node includes: if the first node is the same as the second node, determining the first node as the optimal node; if the first node is different from the second node, selecting the optimal node from the first node and the second node based on a preset convolutional neural network model. The process of selecting the overall optimal node can refer to Figure 2 as shown in Figure 2 The text pool in is the text set to be screened.
[0069] In a preferred implementation manner, after determining the first node as the optimal node if the first node is the same as the second node, it further includes:
[0070] Determine the level composed of the child nodes of the first node as the current level, and loop the steps of determining the first node with the highest node score and the second node with the longest text length in the currently determined level of the DOM tree to determine the optimal nodes corresponding to multiple levels.
[0071] It should be noted that in the process of obtaining the optimal nodes corresponding to multiple levels, it is actually a process of traversal search. That is, starting from the root node of the DOM tree, the child nodes belonging to the current parent node are traversed and searched in turn to screen out the optimal nodes. As a transformable implementation method, in this process, if the text length contained in the subtree is less than 60, the search will be terminated. The nodes of the DOM tree represent a tag and the text under the tag, and a node path can form a complete paragraph. Since the node where the main text is located is the deepest node to be searched, and there is only one path to reach the main text node, compared with the traditional traversal search method in this application, the path search algorithm only needs to compare depth * d times, where depth is the depth of the tree and d is the number of child nodes contained under each parent node on average. Therefore, the path search algorithm can greatly improve the search efficiency.
[0072] In a preferred implementation manner, selecting the optimal node from the first node and the second node based on a preset convolutional neural network model includes: inputting the texts corresponding to the first node and the second node into the preset convolutional neural network model, and outputting the text classification results corresponding to the first node and the second node; and selecting the optimal node from the first node and the second node according to the text classification results corresponding to the first node and the second node.
[0073] In another preferred implementation manner, as Figure 2As shown, according to the text classification results corresponding to the first node and the second node, the optimal node is selected from the first node and the second node, including: if the text classification results corresponding to both the first node and the second node are non-text, the node with the largest number of p tags is selected as the optimal node; if the text classification results corresponding to both the first node and the second node are not non-text, the node with the smallest probability among the non-text tags is selected as the optimal node; if the text classification result corresponding to the first node is non-text, the second node is selected as the optimal node; if the text classification result corresponding to the second node is non-text, the first node is selected as the optimal node. Here, a CNN model is used to classify the text, and the possible tags for each text in the classification result are {N, NP, Y}, where N represents non-text, NP represents possibly containing text, and Y represents text. The specific parameters of the model are shown in the experimental tests. According to the prediction results of the model for the two texts (i.e., the texts corresponding to the above first node and second node), the above four cases can also be expressed as: when the prediction results of both texts are N, based on the number of p tags as the judgment criterion, the node with more p tags is selected as the optimal node; when the prediction results of both texts are not N, the node with the least likelihood of being N (i.e., the smallest probability of N) is selected as the optimal node; when the prediction result of only the text with the highest score (i.e., the first node) is N, the node with the longest text is selected as the optimal node; when the prediction result of only the node with the longest text (the second node) is N, the node with the text having the highest score is selected as the optimal node.
[0074] In some embodiments of the present invention, determining the first node with the highest node score in the current level of the DOM tree includes: calculating the node scores of all nodes in the current level of the DOM tree based on the text density and the symbol density; selecting the node with the highest node score from all nodes in the current level as the first node.
[0075] In some embodiments of the present application, the node scores of all nodes in the current level of the DOM tree are calculated based on the text density and the symbol density, and the formula is:
[0076]
[0077] wherein, td represents the text density of the node, sbd represents the symbol density of the node, p represents the number of p tags, ntd represents the set of text densities in the current level, np represents the set of p tag numbers in the current level, and nsbd represents the set of symbol densities in the current level. The calculation formula for td is:
[0078]
[0079] Among them, t represents the text length under the current node, lt represents the length of the text with links under the current node, tg represents the total number of tags under the current node, and ltg represents the number of tags with links under the current node. The calculation formula of sbd is as follows:
[0080]
[0081] Among them, synb represents the number of symbols under the current node.
[0082] In some embodiments of the present application, as Figure 3 shown, the text set to be screened is screened according to rules to obtain the target text, including: the label predicted by the convolutional neural network model as the text is recorded as the Y label; the longest text corresponding to the Y label is used as the initial solution; the initial solution is screened according to the text length ratio and the linked text ratio to obtain the target text.
[0083] Let the text pool be the set {s1, s2, s3,... s n}, and the longest text predicted by the deep model with the label Y is used as the initial solution, and the boundary is adjusted in a regular way. The calculation method of the text length ratio is as follows:
[0084]
[0085] Among them, t u represents the text length of the upper-level node, and t n represents the text length of the current level. Usually, the number of words of irrelevant information such as the navigation bar is much smaller than the number of words of the text. Therefore, t u < 2t n ; and when t n is part of the text and t u is the text, t u >> 2t n . Therefore, preferably, the present application determines the second-round solution by setting the text length ratio threshold σ = 2. When t u ≤ 2t n , the initial solution is retained. When t u > 2t n , the previous node is selected as the second-round solution. The last-round screening uses the linked text ratio of the current node to determine the solution. The calculation method of the linked text ratio is as follows:
[0086]
[0087] Among them, It n represents the length of the linked text of the current node, and t nIndicates the length of all text (not contradictory to the above, being two dimensions). When Ir is greater than the threshold δ, it can be determined that there is invalid information such as a list page in the current node. The threshold δ is obtained based on experience. Preferably, in this application, it is set to 0.9. The final solution is obtained by screening through the link text ratio. As Figure 3 shown, when Ir ≤ δ, the second-round optimal solution is retained. When Ir > δ, the next-level node is the optimal solution.
[0088] The training process of the deep model adopted in this application is as follows: Using nearly 20,000 node sample data generated from 3,000 web pages, the sample data is divided into a training set and a test set in a ratio of 9:1, and trained using the text-CNN framework. The word dimension used by the model is 128, the vocabulary size is 60,000, the maximum text length is 300, the number of convolutional kernels is 128, the convolutional kernel size is [2, 3, 4], the dropout ratio is 0.5, the learning rate is 0.0001, the batch_size is 64, and the number of training rounds is fixed at 20 rounds. To verify the efficiency and accuracy of the model, this application conducted experiments in the task of extracting the main text of tender web pages. For 3,000 web pages different from the deep model training, the method of this application was used to extract the tender main text, and the evaluation method was the accuracy of extracting the main text, that is, the ratio of the main text extracted by the model to the real main text being exactly matched. The experimental results show that the extraction accuracy of the model is 95%, and the extraction efficiency of the model meets the actual production requirements.
[0089] This application constructs a DOM tree for the tender web page to be extracted, determines the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree, determines the optimal node corresponding to the current level from the first node and the second node, converts the method of extracting the main text of the tender web page into an optimal path search problem, greatly improves the efficiency and reduces the space; and stores the text corresponding to the optimal node into the text set to be screened. The text set to be screened includes the texts of the optimal nodes corresponding to multiple levels. The text set to be screened is screened according to rules to obtain the target main text. Here, the traditional features, deep learning algorithms and rule screening methods are integrated to obtain the target main text, thereby improving the accuracy of extracting the main text.
[0090] In some exemplary embodiments of this application, there is also provided a device for extracting the main text of a tender web page, as Figure 4 shown, the device includes:
[0091] A construction module 401, configured to construct a DOM tree for the tender web page to be extracted;
[0092] A determination module 402, configured to determine the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree;
[0093] A comparison module 403 is configured to determine an optimal node corresponding to the current level from the first node and the second node, and store the text corresponding to the optimal node into a text set to be screened, where the text set to be screened includes texts of optimal nodes corresponding to multiple levels;
[0094] A screening module 404 is configured to perform rule screening on the text set to be screened to obtain a target main text.
[0095] In a specific implementation, the device may further be provided with a search module, which is not specifically limited herein.
[0096] It should also be emphasized that the system provided in the embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0097] Please refer to the following Figure 5 , which shows a schematic diagram of a computer device provided by some embodiments of the present application. As Figure 5 shown, the computer device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the main text extraction method of the bidding web page provided by any of the foregoing embodiments of the present application.
[0098] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0099] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. Among them, the memory 201 is used to store programs. After receiving an execution instruction, the processor 200 executes the program. The method for extracting the main text of the tender web page disclosed in any implementation manner of the embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.
[0100] The processor 200 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in hardware or the instructions in software form in the processor 200. The above-mentioned processor 200 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.
[0101] The embodiment of the present application also provides a computer-readable storage medium corresponding to the method for extracting the main text of the tender web page provided in the foregoing embodiment. Please refer to Figure 6 , Figure 6 The shown computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method for extracting the main text of the tender web page provided in any of the foregoing embodiments.
[0102] In addition, examples of computer-readable storage media can also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0103] The computer-readable storage medium provided in the above embodiments of the present application and the method for allocating quantum key distribution channels in the space-division multiplexing optical network provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0104] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements the steps of the method for extracting the main text of a tender web page provided in any of the foregoing embodiments, including: constructing a DOM tree for the tender web page to be extracted; determining a first node with the highest node score and a second node with the longest text length in the current level of the DOM tree; determining the optimal node corresponding to the current level from the first node and the second node, and storing the text corresponding to the optimal node into a text set to be screened, where the text set to be screened includes the texts of the optimal nodes corresponding to multiple levels; performing rule screening on the text set to be screened to obtain the target main text.
[0105] It should be noted that: The algorithms and displays provided herein are not inherently related to any particular computer, virtual device or other equipment. Various general-purpose devices can also be used in conjunction with the teachings provided herein. Based on the above description, the structures required to construct such devices are obvious. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions of specific languages above are for disclosing the best implementation manners of the present application. In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and technologies are not shown in detail so as not to obscure the understanding of this specification.
[0106] Similarly, it should be understood that, in order to streamline the present application and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than those expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself is considered a separate embodiment of the present application.
[0107] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0108] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program for executing part or all of the methods described herein. The program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0109] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for extracting the main text of a tender web page, characterized in that, The method includes: Constructing a DOM tree for the tender web page to be extracted; Determining a first node with the highest node score and a second node with the longest text length in the current level of the DOM tree, where the node score is calculated based on text density and symbol density; Determining the optimal node corresponding to the current level from the first node and the second node, and storing the text corresponding to the optimal node into a text set to be screened. The text set to be screened includes texts of optimal nodes corresponding to multiple levels. The determination process includes traversing layer by layer from the root node of the DOM tree, performing a path search algorithm comparison on each layer of nodes, and storing the node with the optimal comparison result into the text set to be screened; Performing rule screening on the text set to be screened to obtain the target body text, where the rule screening is based on the text length ratio and link text.
2. The method for extracting the main text of a tender web page according to claim 1, wherein The determining the optimal node corresponding to the current level from the first node and the second node includes: If the first node is the same as the second node, determining the first node as the optimal node; If the first node is different from the second node, selecting the optimal node from the first node and the second node based on a preset convolutional neural network model.
3. The method for extracting the main text of a tender web page according to claim 2, characterized in that, After determining the first node as the optimal node if the first node is the same as the second node, it further includes: Determining the level formed by the child nodes of the first node as the current level, and looping through the step of determining a first node with the highest node score and a second node with the longest text length in the current level of the DOM tree to determine the optimal nodes corresponding to the multiple levels.
4. The method for extracting the main text of a tender web page according to claim 2, characterized in that Selecting the optimal node from the first node and the second node based on a preset convolutional neural network model includes: Inputting the texts corresponding to the first node and the second node into the preset convolutional neural network model, and outputting the text classification results corresponding to the first node and the second node; Selecting the optimal node from the first node and the second node according to the text classification results corresponding to the first node and the second node.
5. The method for extracting the main text of a tender web page according to claim 4, characterized in that, The selecting the optimal node from the first node and the second node according to the text classification results corresponding to the first node and the second node includes: If the text classification results corresponding to both the first node and the second node are non-body text, selecting the node with the largest number of p tags as the optimal node; If the text classification results corresponding to both the first node and the second node are not non-body text, selecting the node with the smallest probability among the non-body text tags as the optimal node; If the text classification result corresponding to the first node is non-body text, selecting the second node as the optimal node; If the text classification result corresponding to the second node is non-body text, selecting the first node as the optimal node.
6. The method for extracting the main text of a tender webpage according to claim 1, characterized in that, Determining a first node with the highest node score in the current level of the DOM tree includes: Calculating the node scores of all nodes in the current level of the DOM tree based on text density and symbol density; Selecting the node with the highest node score from all nodes in the current level as the first node.
7. The method for extracting the main text of a tender web page according to claim 6, characterized in that Calculate the node scores of all nodes at the current level of the DOM tree based on text density and symbol density. The formula is as follows: Among them, td represents the text density of the node, sbd represents the symbol density of the node, p represents the number of p tags, ntd represents the set of text densities at the current level, np represents the set of the number of p tags at the current level, and nsbd represents the set of symbol densities at the current level.
8. The method for extracting the main text of a tender web page according to claim 1, wherein Perform rule screening on the text set to be screened to obtain the target body text, including: Record the label predicted by the convolutional neural network model as the body text as the Y label; take the longest text corresponding to the Y label as the initial solution; Perform rule screening on the initial solution based on the text length ratio and the link text ratio to obtain the target body text.
9. An apparatus for extracting the main text of a tendering web page, characterized in that, The device includes: A construction module for constructing a DOM tree for the tender web page to be extracted; A determination module for determining the first node with the highest node score and the second node with the longest text length in the current level of the DOM tree, where the node score is calculated based on text density and symbol density; A comparison module for determining the optimal node corresponding to the current level from the first node and the second node, and storing the text corresponding to the optimal node in the text set to be screened. The text set to be screened includes the texts of the optimal nodes corresponding to multiple levels. The determination process includes traversing layer by layer from the root node of the DOM tree, comparing each layer of nodes using a path search algorithm, and storing the node with the best comparison result in the text set to be screened; A screening module for performing rule screening on the text set to be screened to obtain the target body text, where the rule screening is based on the text length ratio and the link text.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
11. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Title extraction method and device based on webpage article
CN108268433A
Recognition system and recognition method of non-body text in webpage
WO2014000571A1