Transferable neural architecture for structured data extraction from web documents
Patent Information
- Application Number
- CN202311184681.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2040-01-29
AI Technical Summary
此外,随着基于人工智能的推荐系统和自动化数字助理的引入,在没有个人访问源网站的情况下获得信息已经变得可能
Smart Images

Figure CN117313853B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on January 29, 2020, with application number 202080095203.7 and invention title "Transferable Neural Architecture for Extracting Structured Data from Web Documents". Technical Field
[0002] This disclosure relates to a transferable neural architecture for extracting structured data from Web documents. Background Technology
[0003] Since the advent of the internet, there has been a need for systems and methods to collect, organize, and present information from multiple websites, enabling users to find the content they are looking for effectively and efficiently. This is evident in the continuous development of search engines and algorithms, which allow users to identify and access websites containing information of interest. Furthermore, with the introduction of AI-based recommendation systems and automated digital assistants, obtaining information without personally visiting the source websites has become possible. As the amount of information available on the internet continues to grow, it becomes increasingly difficult for computing systems to effectively parse and catalog relevant information. Summary of the Invention
[0004] This technology relates to systems and methods for efficiently extracting machine-operable structured data from web documents. Using various neural network architectures, this technology is able to create transferable models of information of interest from the raw Hypertext Markup Language (“HTML”) content of a small set of seed websites. These models can then be applied to the raw HTML of other websites to identify similar information of interest without further human input and extract it as structured data for further use by the system and / or other systems. Therefore, this technology is computationally less expensive than systems and methods that rely on visual rendering and can provide improved results tailored to the information of interest. Furthermore, unlike other text-based methods that require building specific extraction procedures for each domain, this technology provides enhanced technical benefits by generating models that can be used across multiple domains, enabling the extraction of machine-operable structured data in a functional form that can be used by other systems.
[0005] In one aspect, this disclosure describes a computer-implemented method for extracting machine-operable data. The method includes: generating a document object model tree for a first page of a first website by one or more processors of a processing system, wherein the document object model tree includes a plurality of nodes, and each of the plurality of nodes includes an XML path (“XPath”) and content; identifying a first node among the plurality of nodes by one or more processors, wherein the content of the first node includes a first sequence of words, and each word in the first sequence includes one or more characters; identifying a second node among the plurality of nodes by one or more processors, wherein the content of the second node includes a second sequence of words, each word in the second sequence including one or more characters, and the second sequence precedes the first sequence on the first page; generating a word-level vector corresponding to each word of the first sequence and the second sequence by one or more processors; and generating a word-level vector corresponding to each word of the first sequence and the second sequence by one or more processors. The sequence comprises: a character-level word vector corresponding to each word in the first sequence; a sequence-level vector generated by one or more processors based on the word-level vector and character-level word vector corresponding to the first sequence; a sequence-level vector generated by one or more processors based on the word-level vector and character-level word vector corresponding to the second sequence; a discrete feature vector corresponding to one or more predefined features in the content of the first node generated by one or more processors; a composite vector of the first node obtained by one or more processors concatenating the sequence-level vector corresponding to the first sequence, the sequence-level vector corresponding to the second sequence, and the discrete feature vector; a node label of the first node generated by one or more processors based on the composite vector of the first node; and structured data extracted from the first node by one or more processors, wherein the structured data associates the content of the first node with the node label of the first node. In some aspects, generating the character-level word vector corresponding to each word in the first and second sequences includes: for each word in the first sequence, encoding the character vector corresponding to each of the one or more characters using a convolutional neural network, and for each word in the second sequence, encoding the character vector corresponding to each of the one or more characters using a convolutional neural network. In some aspects, generating sequence-level vectors based on word-level vectors and character-level word vectors corresponding to the first sequence includes encoding the character-level word vector and word-level vector for each word in the first sequence using a bidirectional long short-term memory neural network. In some aspects, generating sequence-level vectors based on word-level vectors and character-level word vectors corresponding to the second sequence includes encoding the character-level word vector and word-level vector for each word in the second sequence using a bidirectional long short-term memory neural network. In some aspects, generating node labels for the first node based on the composite vector of the first node includes encoding the composite vector of the first node using a multilayer perceptron neural network to obtain the classification of the first node.In some aspects, the node tag of the first node corresponds to one of a plurality of fields of interest. The method may further include: generating a second document object model tree for a second page of a first website by one or more processors, wherein the second document object model tree includes a second plurality of nodes, and each of the second plurality of nodes includes an XPath and content; and extracting a second set of structured data from the second plurality of nodes by one or more processors, wherein the second set of structured data associates the content of each of the second plurality of nodes with the node tag of each of the second plurality of nodes. Furthermore, the method may further include: generating a third document object model tree for a page of the second website by one or more processors, wherein the third document object model tree includes a third plurality of nodes, and each of the third plurality of nodes includes an XPath and content; and extracting a third set of structured data from the third plurality of nodes by one or more processors, wherein the third set of structured data associates the content of each of the third plurality of nodes with the node tag of each of the third plurality of nodes.
[0006] In another aspect, this disclosure describes a computer-implemented method for extracting data, comprising: generating a document object model tree for a first page of a first website by one or more processors of a processing system, wherein the document object model tree includes a first plurality of nodes, and each of the first plurality of nodes includes an XML path (“XPath”) and content; generating a prediction by one or more processors for each of the first plurality of nodes regarding whether the node is associated with one of a plurality of fields of interest; generating a plurality of node pairs from the first plurality of nodes by one or more processors, wherein each of the plurality of node pairs includes a head node and a tail node; generating a composite vector corresponding to each head node and each tail node by one or more processors; generating an XPath vector corresponding to each head node and each tail node by one or more processors; and generating an XPath vector based at least in part on each head node and each tail node by one or more processors. The process involves generating a position vector corresponding to each head node and each tail node relative to the position of at least one other node in a first plurality of nodes; for each node pair, one or more processors concatenate the composite vector, position vector, and XPath vector corresponding to the head and tail nodes of the node pair to obtain a pair-level vector; one or more processors generate a pair label for each node pair based on the pair-level vector; one or more processors generate a node label for the head node of each node pair based on the pair label or a prediction of the head node; one or more processors generate a node label for the tail node of each node pair based on the pair label or a prediction of the tail node; and one or more processors extract structured data from one or more nodes in the first plurality of nodes, wherein the structured data associates the content of each of the one or more nodes with the node label of each of the one or more nodes. In some aspects, generating the XPath vector corresponding to each head node and each tail node includes encoding the XPath of each head node and each tail node using a Long Short-Term Memory (LSTM) neural network. In some aspects, generating a comprehensive vector corresponding to each head node and each tail node includes: for each head node, concatenating a sequence-level vector corresponding to the word sequence in the head node, a sequence-level vector corresponding to the word sequence in the nodes preceding the head node, and a discrete feature vector corresponding to one or more predefined features in the content of the head node; and for each tail node, concatenating a sequence-level vector corresponding to the word sequence in the tail node, a sequence-level vector corresponding to the word sequence in the nodes preceding the tail node, and a discrete feature vector corresponding to one or more predefined features in the content of the tail node. In some aspects, generating pair labels for each node pair based on the pair-level vectors includes: encoding the pair-level vector of each node pair using a multilayer perceptron neural network to obtain a classification for each node pair. In some aspects, node labels correspond to one or an empty identifier among multiple fields of interest.The method may further include: generating a second document object model tree for a second page of a first website by one or more processors, wherein the second document object model tree includes a second plurality of nodes, and each node in the second plurality of nodes includes an XPath and content; and extracting a second set of structured data from the second plurality of nodes by one or more processors, wherein the second set of structured data associates the content of each node in the second plurality of nodes with a node tag of each node in the second plurality of nodes. Furthermore, the method may further include: generating a third document object model tree for a page of the second website by one or more processors, wherein the third document object model tree includes a third plurality of nodes, and each node in the third plurality of nodes includes an XPath and content; and extracting a third set of structured data from the third plurality of nodes by one or more processors, wherein the third set of structured data associates the content of each node in the third plurality of nodes with a node tag of each node in the third plurality of nodes. The method may further include: generating a second document object model tree for a second page of a first website by one or more processors, wherein the second document object model tree includes a second plurality of nodes, and each node in the second plurality of nodes includes an XPath and content; generating node tags for each node in the second plurality of nodes by one or more processors; identifying a class of nodes from the first plurality of nodes and the second plurality of nodes by one or more processors, wherein the node tags of each node in the class are the same; identifying a first XPath as the most common XPath in the class of nodes by one or more processors; and extracting a second set of structured data from each node in the first plurality of nodes and the second plurality of nodes that has the first XPath by one or more processors, the second set of structured data associating the content of the node with the node tags of the node.
[0007] In another aspect, this disclosure describes a processing system for extracting machine-operable data. The processing system includes a memory and one or more processors coupled to the memory and configured to: generate a Document Object Model (DOM) tree for a first page of a first website, wherein the DOM tree includes a plurality of nodes, and each of the plurality of nodes includes an XML path (“XPath”) and content; identify a first node among the plurality of nodes, wherein the content of the first node includes a first sequence of words, and each word in the first sequence includes one or more characters; identify a second node among the plurality of nodes, wherein the content of the second node includes a second sequence of words, each word in the second sequence including one or more characters, and the second sequence precedes the first sequence on the first page; generate a sequence of words corresponding to the first and second sequences. The process involves: generating word-level vectors corresponding to words; generating character-level vectors corresponding to each word in the first and second sequences; generating sequence-level vectors based on the word-level vectors and character-level vectors corresponding to the first sequence; generating sequence-level vectors based on the word-level vectors and character-level vectors corresponding to the second sequence; generating discrete feature vectors corresponding to one or more predefined features in the content of the first node; concatenating the sequence-level vectors corresponding to the first sequence, the sequence-level vectors corresponding to the second sequence, and the discrete feature vectors to obtain a comprehensive vector for the first node; generating a node label for the first node based on the comprehensive vector; and extracting structured data from the first node, wherein the structured data associates the content of the first node with the node label of the first node. In some aspects, the node label of the first node corresponds to one of multiple fields of interest.
[0008] On the other hand, this disclosure describes a processing system for extracting machine-operable data, including a memory and one or more processors, the processors being coupled to the memory and configured to: generate a document object model tree for a first page of a first website, wherein the document object model tree includes a first plurality of nodes, and each of the first plurality of nodes includes an XML path (“XPath”) and content; generate a prediction for each of the first plurality of nodes regarding whether the node is associated with one of a plurality of fields of interest; generate a plurality of node pairs from the first plurality of nodes, wherein each of the plurality of node pairs includes a head node and a tail node; generate a composite vector corresponding to each head node and each tail node; generate an XPath vector corresponding to each head node and each tail node; at least in part based on The position vector corresponding to each head node and each tail node is generated based on the position of each head node and each tail node relative to at least one other node in the first plurality of nodes; for each node pair, the composite vector, position vector, and XPath vector corresponding to the head node and tail node of the node pair are concatenated to obtain a pair-level vector; a pair label is generated for each node pair based on the pair-level vector of the node pair; a node label is generated for the head node of each node pair based on the pair label of the node pair or the prediction of the head node; a node label is generated for the tail node of each node pair based on the pair label of the node pair or the prediction of the tail node; and structured data is extracted from one or more nodes in the first plurality of nodes, wherein the structured data associates the content of each of the one or more nodes with the node label of each of the one or more nodes. In some aspects, the node label corresponds to one or a null value in a plurality of fields of interest.
[0009] In another aspect, this disclosure describes a computer-implemented method comprising: generating a document object model tree for a first page of a website by one or more processors of a processing system, wherein the document object model tree includes a plurality of nodes; generating, by one or more processors, word-level vectors corresponding to each word in a first word sequence associated with a first node in the nodes and a second word sequence associated with a second node in the nodes; generating, by one or more processors, character-level word vectors corresponding to the first and second sequences; generating, by one or more processors, sequence-level vectors based on the word-level vectors and the character-level word vectors; generating, by one or more processors, discrete feature vectors corresponding to one or more features in the content of the first node; generating, by one or more processors, node tags of the first node based on the sequence-level vectors corresponding to the first sequence, the sequence-level vectors corresponding to the second sequence, and the discrete feature vectors; and extracting structured data from the first node by one or more processors, the structured data associating the content of the first node with the node tags of the first node.
[0010] In another aspect, this disclosure describes a processing system configured to extract machine-operable data, comprising: a memory; and one or more processors coupled to the memory, and configured to: generate a document object model tree for a first page of a website, wherein the document object model tree includes a plurality of nodes; generate word-level vectors corresponding to each word in a first word sequence associated with a first node in the nodes and a second word sequence associated with a second node in the nodes; generate character-level word vectors corresponding to the first and second sequences; generate sequence-level vectors based on the word-level vectors and the character-level word vectors; generate discrete feature vectors corresponding to one or more features in the content of the first node; generate node tags for the first node based on the sequence-level vectors corresponding to the first sequence, the sequence-level vectors corresponding to the second sequence, and the discrete feature vectors; and extract structured data from the first node, the structured data associating the content of the first node with the node tags of the first node. Attached Figure Description
[0011] Figure 1 It is a functional diagram of an example system based on various aspects of this disclosure.
[0012] Figure 2 It is a diagram showing how a portion of HTML can be represented as a DOM tree.
[0013] Figure 3 This is a flowchart of exemplary methods according to various aspects of this disclosure.
[0014] Figure 4 This is a flowchart of exemplary methods according to various aspects of this disclosure.
[0015] Figure 5 This is a flowchart of exemplary methods according to various aspects of this disclosure.
[0016] Figure 6 This is a diagram illustrating how exemplary phrases can be processed according to various aspects of this disclosure.
[0017] Figure 7 This is a flowchart of exemplary methods according to various aspects of this disclosure.
[0018] Figure 8 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure.
[0019] Figure 9 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure.
[0020] Figure 10 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure.
[0021] Figure 11 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure.
[0022] Figure 12 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure.
[0023] Figure 13 This is a flowchart illustrating exemplary methods based on various aspects of this disclosure. Detailed Implementation
[0024] This technology will now be described with respect to the following exemplary systems and methods.
[0025] Example System
[0026] Figure 1 An arrangement 100 having an exemplary processing system 102 for performing the methods described herein is schematically illustrated. Processing system 102 includes one or more processors 104 and memory 106 for storing instructions and data. Additionally, the one or more processors 104 may include various modules described herein, and the instructions and data may include various neural networks described herein. Processing system 102 is shown communicating with various websites (including websites 110 and 118) via one or more networks 108. Exemplary websites 110 and 118 each include one or more servers 112a-112n and 120a-n, respectively. Each of servers 112a-112n and 120a-n may have one or more processors (e.g., 114 and 122) and associated memory (e.g., 116 and 124) storing instructions and data (including HTML of one or more web pages). However, various other topologies are also possible. For example, processing system 102 may not communicate directly with the websites and may instead process a stored version of the HTML of the website to be processed.
[0027] Processing system 102 can be implemented on different types of computing devices, such as any type of general-purpose computing device, server, or a combination thereof, and may also include other components typically found in general-purpose computing devices or servers. Memory 106 stores information accessible by one or more processors 104, including instructions and data that can be executed or otherwise used by processor 104. Memory can be any non-transitory type capable of storing information accessible by processor 104. For example, memory can include non-transitory media such as hard disk drives, memory cards, optical discs, solid-state storage, magnetic tape storage, etc. Computing devices suitable for the role described herein may include different combinations of the foregoing, whereby different portions of instructions and data are stored on different types of media.
[0028] In all cases, the computing device described herein may also include any other components typically used in conjunction with a computing device such as a user interface subsystem. The user interface subsystem may include one or more user inputs (e.g., a mouse, keyboard, touchscreen, and / or microphone) and one or more electronic displays (e.g., a monitor with a screen or any other electrical device operable to display information). Output devices other than electronic displays (such as speakers, lights, and vibrational, pulsating, or haptic elements) may also be included in the computing device described herein.
[0029] The one or more processors included in each computing device can be any conventional processor, such as a commercially available central processing unit (“CPU”), tensor processing unit (“TPU”), etc. Alternatively, the one or more processors can be special-purpose devices, such as ASICs or other hardware-based processors. Each processor can have multiple cores capable of operating in parallel. The processor, memory, and other components of a single computing device can be stored within a single physical enclosure or distributed across two or more enclosures. Similarly, the memory of a computing device can include hard drives or other storage media located in an enclosure different from the processor enclosure (such as in an external database or networked storage device). Therefore, references to processors or computing devices will be understood to include a collection of processors or computing devices or memory that may or may not operate in parallel, as well as one or more servers in a load-balanced server cluster or cloud-based system.
[0030] The computing device described herein can store instructions that can be executed directly by a processor (such as machine code) or indirectly (such as scripts). The computing device can also store data that can be retrieved, stored, or modified by one or more processors according to the instructions. Instructions can be stored as computing device code on a computing device-readable medium. In this regard, the terms "instruction" and "program" are used interchangeably herein. Instructions can also be stored in an object code format for direct processor processing, or in any other computing device language, including scripts or collections of independent source code modules that are interpreted on demand or pre-compiled. As an example, the programming language could be C#, C++, JAVA, or another computer programming language. Similarly, any component of the instructions or program can be implemented in a computer scripting language, such as JavaScript, PHP, ASP, or any other computer scripting language. Furthermore, any of these components can be implemented using a combination of computer programming languages and computer scripting languages.
[0031] Example Method
[0032] In addition to the systems described above and in the accompanying figures, various operations will now be described. In this regard, there are multiple ways in which the processing system 102 can be configured to extract structured data from websites. For example, the processing system 102 can be configured to use a site-specific extraction procedure or “wrapper” for each website from which data is to be extracted. However, site-specific methods typically require human analysis of the site and the creation of a wrapper to be used by the extraction procedure, or require that the website’s pages be annotated well enough so that the extraction procedure can accurately identify pre-selected fields of interest without human input. In either case, a wrapper created for one site cannot be transferred to a different site. Fields of interest can be any category of information selected for extraction. For example, for a website related to automobiles, fields of interest could include model name, vehicle type, mileage, engine size, engine power, engine torque, etc.
[0033] In other cases, it is possible to train neural networks on a set of rendered web pages to identify information of interest using various visual cues. However, while visual rendering methods can generate models that allow for the identification and extraction of fields of interest from other websites, they require careful feature engineering using domain-specific knowledge to generate the models and are computationally expensive due to rendering.
[0034] Given these drawbacks, this technique provides a neural network architecture capable of creating transferable extraction models using text from a set of seed websites, with minimal or no human input and without the need for rendering. These extraction models can then be used to identify and extract information of interest from the text of additional websites without requiring any webpage rendering.
[0035] In this regard, in an exemplary method according to various aspects of the present technology, the processing system first applies a node-level module to a selected set of seed websites. Seed websites can be selected based on various attributes. For example, some websites will already include annotations identifying various fields of interest. In this regard, on an exemplary automotive website, each vehicle's page may have a table with rows labeled "Model" followed by the model name, labeled "Engine" followed by the engine size, labeled "Gasoline Mileage" followed by the gasoline mileage, and so on. Sites with one or more annotations associated with pre-selected fields of interest may be helpful as seed websites, as they allow the neural network to generate models that are more accurate in identifying fields of interest in other websites with fewer annotations. The node-level module parses the raw HTML of each page from each seed website into a Document Object Model ("DOM") tree. This transforms each page into a tree structure where each branch ends at a node, and each node includes an XML path ("XPath") and its associated HTML content. For example, Figure 2 Diagram 200 illustrates how a portion of HTML 202 can be represented as a DOM tree 204. While the nodes of the DOM tree 204 are... Figure 2 They are shown as empty circles, but in reality they will contain the HTML content associated with each node.
[0036] The node-level module then identifies all nodes containing text and filters the list of all such text nodes to remove those that are unlikely to convey information of interest. This can be done, for example, by collecting all possible XPaths (node identifiers) for text nodes in a given website, ranking the XPaths according to the number of distinct text values associated with each XPath, and identifying a subset of those with two or more distinct values as nodes of interest. Figure 3 This includes a flowchart 300 illustrating this exemplary filtering method. In this regard, at step 302, the node-level module parses the raw HTML of the webpage into a DOM tree. At step 304, the node-level module identifies all nodes in the DOM tree that contain text. At step 306, the node-level module identifies all XPaths associated with the text nodes in the DOM tree. At step 308, the node-level module ranks all XPaths based on how many distinct text values are associated with each such XPath. In step 310, the node-level module identifies the top N XPaths with two or more distinct text values as “nodes of interest.” Therefore, in some examples, the node-level module may rank XPaths according to the number of distinct values associated with them and select the top 500 XPaths (or more, or fewer) with at least two distinct values. Filtering in this manner will remove most nodes that share common values across multiple pages and are therefore more likely to represent common text on the pages of the website, such as the website name, navigation text, headers, footers, copyright information, etc.
[0037] The node-level module then encodes the filtered set of text nodes (“nodes of interest”) using the text of each node, the text preceding each node, and one or more discrete features (e.g., content in the original HTML that might help identify fields of interest). The following will discuss… Figure 4-9 In further detail, each of these encoding processes utilizes a different neural network.
[0038] In this regard, such as Figure 4 As shown in step 402 of method 400, when encoding based on the text of each node, the node-level module decomposes the text of each node into {w1, w2, ..., w...} |n| The word sequence W is composed of {w1, w2, ..., w}.|n|} can be the original text of the node, or it can be the result of lexical analysis of the original text, such as by tokenizing and lexicalizing the original text using a Natural Language Toolkit (“NLTK”). Therefore, for each node, each word w i Elements that can be represented as W according to Equation 1 below. As used herein, “words” do not need to consist of letters and are therefore able to include text consisting of numbers and / or symbols, such as “$1,000”.
[0039] w i ∈W(1)
[0040] As shown in step 404, the node-level module will also process each word w i Decomposed into {c1,c2,...,c |w|} i The character sequence C is formed. Therefore, for a given word w in a node... i Each character c j Elements that can be represented as C according to Equation 2.
[0041] c j ∈C(2)
[0042] As shown in step 406, the character embedding lookup table E c It is also initialized. Step 406 can be performed before steps 402 and / or 404. Character embedding lookup table E c According to the definition in Equation 3, where dim c is a hyperparameter representing the dimension of the character embedding vector, and It is the symbol representing all real numbers. Therefore, the character embedding lookup table E c It is a shape of |C|x dim c A matrix, where each element is a real number. E c The character embedding vectors are randomly initialized and then updated via backpropagation during model training. The dimension of the character embedding vectors can be any suitable number, such as 100 (or more or less).
[0043]
[0044] As shown in step 408, for each word w i Using character embedding lookup table E c For each character c j Generate character embedding vectors. Next, in step 410, a convolutional neural network (“CNN”) is used to process each word w. i The entire sequence of character embedding vectors is encoded, and then pooled to create words with the same embedding vectors.i The corresponding character-level word vector c i Therefore, the character-level word vector c i This can be represented by Equation 4 below. These steps are also... Figure 6 The diagram illustrates that the exemplary phrase "city25hwy 32" is processed to create individual character embedding vectors for each word 602a-602d. Then, each set of character embedding vectors for each word is fed into a CNN 606 and pooled to create corresponding character-level word vectors 608a-608d. The CNN can be implemented with any suitable parameters. For example, the CNN can have a kernel size of 3 (or larger or smaller), a filter size of 50 (or larger or smaller), and can apply max pooling to select the maximum value from each row of the resulting matrix, thus reducing each row to a single value.
[0045] c i =CNN({c1,c2,…,c |w|}) (4)
[0046] Additionally, as shown in step 412, each word w is also... i Initialize word-level vector lookup table E w Here, step 412 may also occur before any or all of steps 402-410. Word-level vector lookup table E w Defined according to Equation 5 below, where dim w E is a hyperparameter representing the dimension of word-level vectors. w The word-level vectors in the table can be generated from various known algorithms, such as Stanford's GloVe. Therefore, the word-level vector lookup table E... w It is a shape of |W|xdim w The matrix is a set of numbers, and each element in the matrix is a real number. The dimension of the character embedding vector can be any suitable number, such as 100 (or more or less).
[0047]
[0048] As shown in step 414, the word-level vector lookup table E is used. w Generate each word w i Word-level vector w i Then, as shown in step 416, for each word w i Word-level vectors w i Compared to the character-level word vectors c created by CNN i Cascade, so that each word in each node is w i Create cascaded word vectors t iThis is shown in Equation 6 below, where [·⊙·] denotes a cascading operation. These steps are also... Figure 6 The diagram illustrates that each word 604a-604d in the phrase "city 25hwy 32" is processed to create a corresponding word-level vector 610a-610d. These word-level vectors 610a-610d are then concatenated with their associated character-level word vectors 608a-608d to form concatenated word vectors 612a-612d.
[0049] t i =[w i ⊙c i (6)
[0050] As a result of the above, for a set of all words W in a given node, there will exist a set consisting of {t1,t2,…,t…} |n| The set of concatenated word vectors. Next, as... Figure 5 and Figure 6 As shown, a Long Short-Term Memory (“LSTM”) neural network is used to perform context encoding on the set of concatenated word vectors from both forward and backward directions. In other words, the LSTM network will encode the set {t1,t2,…,t…} from both forward and backward directions. |n|} and the inverse set {t |n| ,t |n|-1 Encoding is performed on t1, ..., t2. These processes are respectively performed on t1, ..., t2. Figure 5 Steps 502 and 504 of process 500 are described. The forward and backward LSTM encoding steps are also described separately by... Figure 6 The dashed lines 614 and 616 are graphically illustrated. Because the LSTM network encodes the set in both directions, it can also be called a bidirectional long short-term memory neural network. Although the LSTM network will have the same structure in both directions, the parameters will be different for the forward and backward encoding operations. The LSTM network can take any suitable number of units, such as 100 (or more or less). The results of the forward and backward LSTM encoding operations are then subjected to mean pooling to reach the final representation n of the text at that node. node_text ,like Figure 5 Step 506 is shown. This can be represented as shown in Equation 7, where AVG[·⊙·] represents the average or mean pooling operation, and where LSTM f and LSTM b These represent LSTM operations in the forward and backward directions, respectively. This is also... Figure 6 The diagram illustrates that the outputs of the forward and backward LSTM encoding operations are represented by dashed lines 618 and 620, the mean pooling of those outputs is shown in element 622, and the final vector representation of the text is shown in element 624.
[0051] n node_text =AVG[LSTM f ({t1,t2,…,t |n| LSTM b ({t |n| ,t |n|-1 ,…,t1})] (7)
[0052] Encoding both words and characters in the text of each node, as described above, enables node-level modules to recognize patterns shared across nodes, even when the individual words in a given node may be unknown (e.g., misspellings, abbreviations, etc.), or when the node's text includes numbers or special characters. For example, in the context of a website about cars, nodes might contain patterns such as... Figure 6 The text shown is “city 25hwy 32”, but nodes from another page on the same website could contain similar text for another car, such as “city 25hwy 28”. By simply tracking words, the node-level module could only determine that these nodes share the words “city” and “hwy,” the latter of which might not even be recognized as a word because it is merely an abbreviation. However, by combining the results of character-level CNN and word-level LSTM operations as described above, the node-level module is able to identify that these nodes actually share the pattern “city##hwy##”. Importantly, the node-level module is able to do this without requiring human input.
[0053] As described above, and respectively as Figure 7 and Figure 8 As shown in flowcharts 700 and 800, the node-level module also encodes the text preceding each node of interest. This encoding can be performed on a fixed amount of preceding text. The node-level module is related to the above... Figure 4-6 The preceding text is processed in the same way as described, thus producing the second vector n. prev_text . Figure 4 and Figure 5 The steps do not need to be in Figure 7 and Figure 8 The process occurs before the previous step. Conversely, node-level modules can process previous text before, after, or simultaneously with their processing of node text.
[0054] Therefore, as Figure 7 As shown in step 702, the node-level module decomposes the text preceding the node of interest into a sequence of X words. X can be any number, such as 10 words (or more or fewer words). Additionally, as mentioned above regarding... Figure 4As described in step 402, the sequence of X words can be the original text preceding the node of interest, or it can be the result of subjecting the preceding text to lexical analysis, such as tokenization and lemmatization using the NLTK toolkit. At step 704, the node-level module is aligned with the above regarding... Figure 4 Step 404 describes the same method used to decompose each of the X words into a sequence of characters. At step 706, the node-level module, in accordance with the above... Figure 4 The character embedding lookup table is initialized in the same manner as described in step 406. (And...) Figure 4 The steps are the same. Figure 7 Step 706 can occur before steps 702 and / or 704. Furthermore, in this respect, node-level modules can use the same character embedding lookup table for... Figure 4 and Figure 7 In this process, steps 406 and 706 will both describe initializing a single instance of the character embedding lookup table. At step 708, the node-level module uses the character embedding lookup table to encode the characters of each of the X words, in accordance with the above... Figure 4 The corresponding character embedding vector is created in the same manner as described in step 408. At step 710, for each of the X words, the node-level module uses a CNN to encode the corresponding sequence of character embedding vectors, and then... (The sentence is incomplete and requires more context to translate accurately.) Figure 4 Step 410 describes pooling them in the same way to create character-level word vectors for each word. Step 710 can use a combination of... Figure 4 Step 410 uses the same CNN, or a separate CNN can be used. At step 712, the node-level modules are configured as described above. Figure 4 The word-level vector lookup table is initialized in the same manner as described in step 412. Here, Figure 7 Step 712 can also be done in Figure 7 This occurs before any or all of steps 702-710. Furthermore, in this respect, the node-level module can use the same word-level vector lookup table. Figure 4 and 7 In this process, steps 412 and 712 will both describe initializing a single instance of the word-level vector lookup table. At step 714, for each of the X words, the node-level module encodes the word using the word-level vector lookup table to correspond with the above... Figure 4 Step 414 describes the creation of the corresponding word-level vectors in the same manner. In step 716, for each of the X words, the node-level module creates a vector in the same way as described above. Figure 4 Step 416 describes the same method of concatenating the corresponding word-level vectors and character-level word vectors to create concatenated word representations.
[0055] Similarly, in Figure 8 In steps 802 and 804, for a sequence of X words, the node-level modules respectively use the above-mentioned... Figure 5 Steps 502 and 504 describe the same approach, using an LSTM network to represent the corresponding concatenated words in both the forward and backward directions (in... Figure 7 The LSTM network created in step 716 is used for encoding. Here, steps 802 and 804 can also use the same LSTM network used in conjunction with steps 502 and 504, or a separate LSTM network can be used. Finally, at step 806, the results of the forward and backward LSTM encoding operations undergo mean pooling to correlate with the values above regarding the encoding process. Figure 5 In step 506, n is generated node_text The final representation of the previous text that reaches the node of interest in the same way as described. prev_text .
[0056] Encoding the preceding text of each node, as described above, can further help distinguish nodes with similar content. For example, in a website about cars, node-level modules can be programmed to identify the gasoline mileage value on each page. Thus, a given page might include a first node with the text "25" and a second node with the text "32". The text of these two nodes alone may not contain enough information to determine whether either one represents a gasoline mileage value. However, in many cases, the text preceding these nodes will contain descriptive terms such as "gasoline mileage," "fuel economy," "miles per gallon," "highway mileage," or some other text that strengthens or weakens that inference.
[0057] As mentioned above, the node-level module also examines the text of each node against a pre-selected set of discrete features, such as Figure 9 The process is shown in step 900. This yields the third vector n. dis_feat .here, Figure 4-8 The steps are also unnecessary Figure 9 The process occurs before the previous step. Conversely, node-level modules can examine discrete features before, after, or simultaneously with their processing of node text and / or preceding text.
[0058] Therefore, as Figure 9 As shown in step 902, the node-level module initializes a discrete feature lookup table E containing a preselected set of discrete features of interest. d These discrete features can be anything in the original HTML that is identified as helpful in identifying fields of interest. For example, in many cases, the leaf tag type of a node (e.g., <h1>, , , This will help categorize information on the page. In this respect, <h1>Nodes are more likely to include key information, such as the model name of the vehicle being displayed on the page. Similarly, known algorithms, such as the string type checker in the NLTK toolkit, can be used to determine whether the text of a given node includes information of a selected type that might be helpful, such as dates, postal codes, or URL links. These and any other discrete features deemed to be of interest can be included in the discrete feature lookup table E. d Therefore, according to Equation 8 below, the discrete feature lookup table E is defined. d Where D is the set of all identified discrete features, and dim d This is a hyperparameter representing the dimension of the discrete feature vector. The dimension of the discrete feature vector can be any suitable number, such as 30 (or more or less).
[0059]
[0060] exist Figure 9 In step 904, the node-level module then generates a vector d, where each of the preselected discrete features present for a given node is represented as a non-negative integer. For example, if the preselected discrete features of a given set of websites are {gasoline mileage, date, postal code}, and the node of interest has two gasoline mileage values, one date, and no postal code, then the vector d for that node will have the values {2, 1, 0}. Therefore, the vector d is defined according to Equation 9 below, where, It is the symbol for all non-negative integers.
[0061]
[0062] exist Figure 9 In step 906, for each node of interest, the node-level module uses matrix multiplication according to Equation 10 below to combine the vector d representing the discrete features with the discrete feature lookup table E. d Multiplication. This yields a single vector n. dis_feat It is the final representation of the discrete features existing in the nodes of interest.
[0063] n dis_feat =dE d (10)
[0064] Once the three encoding processes have been executed, the node-level module uses the n obtained from each node. node_text n prev _text and n dis_feat Vectors are used to predict whether a node corresponds to one of a predefined set of fields of interest. For example, for an automotive website, fields of interest might include model name, vehicle type, engine, and mileage, and the node-level module will use the final vector generated for each node to predict whether the node corresponds to any of those fields of interest. If it does, the node will be labeled according to its corresponding field of interest. If not, the node will be labeled with an empty identifier such as "none" or "empty". See below for reference. Figure 10 The process is further elaborated in detail in step 1000.
[0065] In this regard, Figure 10 At step 1002, the node-level module will display the final representation (n) of the text of the node of interest. node_text The final representation of the text preceding the node of interest (n) prev_text ) and the final representation of the discrete features existing in the nodes of interest (n dis_feat The nodes are cascaded to create a single vector n, which is a comprehensive representation of each node. Therefore, vector n is described according to Equation 11 below.
[0066] n = [n node_text ⊙n prev_text ⊙n dis_feat (11)
[0067] exist Figure 10 At step 1004, the node-level module connects the vector n to a multilayer perceptron (MLP) neural network for multi-class classification via a Softmax function. As shown in step 1006, based on multi-class classification, the node-level module predicts the label l for each node of interest. Since the label l can be any of K predefined fields, or an empty identifier (e.g., "none", "null", etc.), l has K+1 possible values. Therefore, Softmax normalization can be described according to equations 12 and 13 below, where the label l will be the set {f1, ..., f...} K The MLP network can be implemented with any suitable parameters. For example, the MLP network can be a single-layer dense neural network with K+1 nodes, such that the output h is a vector of length K+1.
[0068]
[0069]
[0070] As described above, node-level modules can predict the label l for each node of interest based solely on the node's text, its preceding text, and selected discrete features. However, because the predictions made by node-level modules are all performed discretely for individual nodes of interest, they do not consider what predictions have already been made for other nodes. In some cases, this can lead to node-level modules assigning the same label to multiple nodes on the page, failing to assign other labels to any nodes on the page. Therefore, to further improve the prediction for each node of interest, this technique can also employ a second-stage module that processes node pairs through a relational neural network, such as... Figure 11 The process is shown in step 1100.
[0071] In this respect, the second-stage module can process every possible pair of nodes or a subset thereof on a given webpage, in which case the processing will begin from... Figure 11 Step 1110 begins. However, this may not be feasible in all cases. For example, if the node-level module identifies and encodes 300 nodes on a page, the second-stage module would have 89,700 node pairs to process (i.e., 300 × 299, since the order of the head and tail nodes is important in this context), which may be computationally too expensive. Therefore, in some aspects of this technique, the second-stage module can alternatively divide the segments of interest into two groups, such as... Figure 11 Steps 1102 and 1104 are shown below. Therefore, at step 1102, the second-stage module identifies all fields for which the node-level module predicts at least one node; these will be referred to hereafter as "determined fields." Similarly, at step 1104, the second-stage module identifies all fields for which the node-level module cannot predict any node; these will be referred to hereafter as "uncertain fields." Then, at step 1108, the second-stage module creates all possible node pairs from the subsequent set of nodes. For each determined field, the second-stage module uses the node predicted for that field. For each uncertain field, as shown in step 1106, the second-stage module uses the h generated by the node-level module for that field according to equations 12 and 13 above. i The fraction is used to utilize the first m nodes (e.g., m can be between 5 and 20, or more or less). This will result in three types of node pairs. Nodes for each defined field will be paired with nodes for all other defined fields. Therefore, if there are T defined fields, there will be T(T-1) node pairs consisting entirely of nodes for two defined fields. Furthermore, nodes for each defined field will be paired with the first m nodes identified for each uncertain field. Therefore, if there are K fields in total, this results in an additional 2T(m(KT)) such node pairs, since the order of the head and tail nodes is important in this context. Finally, the first m nodes identified for each uncertain field will be paired with the first m nodes identified for all other uncertain fields. This results in an additional m 2 There are (KT)(KT-1) such node pairs. Therefore, as Figure 11 The total number of node pairs generated as a result of step 1108 can be represented by the following equation 14.
[0072] node_pairs=(T(T-1)+2T(m(KT))+m 2 (KT)(KT-1)) (14)
[0073] Then, the second-stage module processes each node pair (n) through a relational neural network. head ,n tail ), so as to predict the label (l) head ,l tail To this end, the second-stage module handles node pairs in two ways based on the assumption that two nodes that are closer to each other are more likely to be similar to each other.
[0074] In one scenario, as shown in step 1110, the second-stage module processes each node pair based on the XPath of the head and tail nodes. In this regard, each XPath can be viewed as a sequence of HTML tags, such as "", "", etc. ”、" "and" The second-stage module maintains an embedding matrix for all possible HTML tags, where each tag is represented as a vector. Then, an LSTM network (which may be a different network than the LSTM network used by the node-level modules) uses this matrix to encode each node pair based on their XPath, as shown in Equation 15 below. This yields vectors for the head and tail nodes, respectively. and LSTM networks can use any suitable number of units, such as 100 (or more or fewer).
[0075] n xpath =LSTM([tag1,tag2,...]) (15)
[0076] In another scenario, as shown in steps 1112 and 1114, the second-stage module processes each node pair based on its position relative to other nodes on the original HTML page. In this regard, as shown in step 1112, each node of interest on the page is assigned a position value based on its order relative to the total number of nodes of interest. For example, for a page with 500 nodes of interest, the fifth node could be assigned the value 5. As another example, a scaling value such as 5 / 500 = 0.01 could be assigned to the fifth node. As further shown in step 1112, the second-stage module then initializes a position embedding lookup table E indexed according to each position value. pos , where each position value is associated with a position embedding vector. E pos The location embedding vectors are randomly initialized and then updated via backpropagation during model training. Then, as shown in step 1114, the second-stage module uses the location embedding lookup table E. pos To obtain the vectors containing the positions of the head and tail nodes of each node pair. and
[0077] In addition to the above, the second-stage module also utilizes the comprehensive node vector n generated by the node-level module for each head and tail node, that is, the n vectors generated according to Equation 11 above. Therefore, in step 1116, the second-stage module will synthesize the node vector n head and n tail with vector and (from Equation 15) and and The vectors are cascaded to reach a single synthesized node pair r, as shown in Equation 16 below.
[0078]
[0079] like Figure 12 As shown in step 1202 of process 1200, for each node pair, the second-stage module then connects the synthesized node pair vector r to an MLP network for multi-class classification (which can be a different network from the MLP network used by the node-level module) via a Softmax function. This MLP network can be implemented with any suitable parameters. For example, the MLP network can be a single-layer dense neural network with four nodes, such that the output is a 1×4 vector. The vector output by the MLP is then normalized using the Softmax function. Based on this classification, in step 1204, the second-stage module assigns a normalized label to each node pair. The normalized label is selected from the set {"none-none", "none-value", "value-none", "value-value"}.
[0080] As shown in step 1206, for each determined field, the second-stage module uses the nodes predicted by the first-stage module as the final prediction (or multiple predictions) for that field. As shown in step 1208, for each uncertain field, the second-stage module determines whether any of the m candidate nodes initially identified as that field has been classified as a "value" in any node pair that already includes them. If so, at step 1210, the second-stage module uses that field as the final prediction for that node. For example, for field F and candidate node y (which is one of the m candidate nodes initially identified as field F), there may be four node pairs involving node y. If node y receives a "none" label in three of these pairs and a "value" label in one of these pairs, then the final prediction for node y will be that it corresponds to field F. Finally, as shown in step 1212, based on these final predictions, the processing system extracts structured data from each identified node on each page of each seed website. Importantly, this technique allows the processing system to extract the web data in a structured form that preserves the association between the data and its predicted fields of interest. For example, the extracted data could include data about a car with four fields of interest, where the data is associated with a tag for each of these fields, such as {Model Name|328xi, Vehicle Type|Coupe, Engine|3.0L Inline 6-cylinder, Gasoline Mileage|17 / 25mpg}. This produces functional data. That is, the data is a machine-executable, structured form that allows it to be used to control the operation of processing systems and / or other systems for various purposes. For example, structured data can be used to improve search results in search engines or to create databases from different data sources. In another example, structured data from a website or HTML-based email or message could be used by an automated assistant to provide answers to questions or automatically add events to a user's calendar. Of course, these examples are not intended to be limiting.
[0081] In addition to the above, such as Figure 13 As shown in flowchart 1300, the second-stage module can also utilize additional heuristics to improve the model's predictions. In this regard, for some websites, nodes associated with a specific field can have a relatively small amount of XPath across various pages. Therefore, as... Figure 13 As shown in step 1302, after the second-stage module has generated its first prediction set, for each field of interest f k It can rank how frequently an XPath field is predicted across all pages. Then, as shown in step 1304, for each field of interest f k The second-stage module is able to extract data from the most frequently predicted XPath of this field. This extraction is possible in addition to extracting data from any node(s) indicated by the first prediction set of this field, as described above. Figure 12 As stated above.
[0082] Finally, once the processing system has generated final predictions for all pages across all seed websites as described above, it can perform the same processing steps to generate final predictions for an additional set of non-seed websites. These additional final predictions can then be used to extract further structured data from those non-seed websites in the same manner as described above. In this respect, as a result of first exposing the neural network to seed websites with data on more annotations, organization, current and / or complete fields of interest, the model built by the neural network will be more accurately able to identify fields of interest in non-seed websites that may not have the same annotations, organization, current and / or complete data. Therefore, this technique enables the generation of models with little or no human input, which can then be transferred to allow for the efficient extraction of structured functional data across multiple domains.
[0083] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive, but can be implemented in various combinations to achieve unique advantages. Because these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the exemplary systems and methods should be considered illustrative rather than limiting of the subject matter defined by the claims. Furthermore, the provision of examples described herein and terms such as "such as," "comprising," "including," etc., should not be construed as limiting the subject matter of the claims to the specific examples; rather, these examples are intended only to illustrate some embodiments among many possible implementations. Moreover, the same reference numerals in different figures can identify the same or similar elements. < / h1> < / h1>
Claims
1. A computer-implemented method, comprising: One or more processors of the processing system generate a document object model tree for the first page of the website, wherein the document object model tree includes multiple nodes; The word-level vector corresponding to each word in the first sequence of words associated with the first node in the node and the second sequence of words associated with the second node in the node, generated by one or more processors; One or more processors generate character-level word vectors corresponding to the first and second sequences; One or more processors generate sequence-level vectors based on word-level vectors and character-level word vectors; One or more processors generate discrete feature vectors corresponding to one or more features in the content of the first node; The node label of the first node is generated by one or more processors based on the sequence-level vector corresponding to the first sequence, the sequence-level vector corresponding to the second sequence, and the discrete feature vector; and One or more processors extract structured data from the first node, which associates the content of the first node with the node tag of the first node.
2. The method according to claim 1, wherein, Each of the plurality of nodes includes an XML path and content.
3. The method according to claim 1, wherein, Generating character-level word vectors corresponding to the first sequence includes: For each word in the first sequence, a convolutional neural network is used to encode the character vector corresponding to each character.
4. The method according to claim 3, wherein, Generating character-level word vectors corresponding to the second sequence includes: For each word in the second sequence, a convolutional neural network is used to encode the character vector corresponding to each character.
5. The method according to claim 1, wherein, Generating sequence-level vectors based on word-level vectors and character-level word vectors includes: A bidirectional long short-term memory neural network is used to encode the character-level word vector and word-level vector for each word in the first sequence.
6. The method according to claim 5, wherein, Generating sequence-level vectors based on word-level vectors and character-level word-level vectors also includes: A bidirectional long short-term memory neural network is used to encode the character-level word vector and word-level vector for each word in the second sequence.
7. The method according to claim 1, wherein, The node markers that generate the first node include: The first node's composite vector is encoded using a multilayer perceptron neural network to obtain the classification of the first node.
8. The method according to claim 7, wherein, A composite vector is obtained by concatenating the sequence-level vectors corresponding to the first sequence and the sequence-level vectors corresponding to the second sequence.
9. The method according to claim 8, wherein, The concatenation also includes concatenating discrete feature vectors to sequence-level vectors corresponding to the first sequence and sequence-level vectors corresponding to the second sequence.
10. The method according to claim 1, wherein, On the first page of the website, the second sequence precedes the first sequence.
11. A processing system configured to extract machine action data, comprising: Memory; and One or more processors are coupled to the memory and configured to: Generate a document object model tree for the first page of the website, wherein the document object model tree includes multiple nodes; Generate word-level vectors for each word in the first sequence of words associated with the first node in the node and the second sequence of words associated with the second node in the node; Generate character-level word vectors corresponding to the first and second sequences; Generate sequence-level vectors based on word-level vectors and character-level word vectors; Generate discrete feature vectors corresponding to one or more features in the content of the first node; The node label of the first node is generated based on the sequence-level vector corresponding to the first sequence, the sequence-level vector corresponding to the second sequence, and the discrete feature vector; and Structured data is extracted from the first node, which associates the content of the first node with the node tag of the first node.
12. The processing system according to claim 11, wherein, Each of the plurality of nodes includes an XML path and content.
13. The processing system according to claim 11, wherein, Generating character-level word vectors corresponding to the first sequence includes: For each word in the first sequence, a convolutional neural network is used to encode the character vector corresponding to each character.
14. The processing system according to claim 13, wherein, Generating character-level word vectors corresponding to the second sequence includes: For each word in the second sequence, a convolutional neural network is used to encode the character vector corresponding to each character.
15. The processing system according to claim 11, wherein, Generating sequence-level vectors based on word-level vectors and character-level word vectors includes: A bidirectional long short-term memory neural network is used to encode the character-level word vector and word-level vector for each word in the first sequence.
16. The processing system according to claim 15, wherein, Generating sequence-level vectors based on word-level vectors and character-level word-level vectors also includes: A bidirectional long short-term memory neural network is used to encode the character-level word vector and word-level vector for each word in the second sequence.
17. The processing system according to claim 11, wherein, The node markers that generate the first node include: The first node's composite vector is encoded using a multilayer perceptron neural network to obtain the classification of the first node.
18. The processing system according to claim 17, wherein, A composite vector is obtained by concatenating the sequence-level vectors corresponding to the first sequence and the sequence-level vectors corresponding to the second sequence.
19. The processing system according to claim 18, wherein, The concatenation also includes concatenating discrete feature vectors to sequence-level vectors corresponding to the first sequence and sequence-level vectors corresponding to the second sequence.
20. The processing system according to claim 11, wherein, On the first page of the website, the second sequence precedes the first sequence.
Citation Information
Patent Citations
Neural network-based scholarship user portrait information extraction method and model
CN109657135A
Handling tabular information
EP1550961A1