Method, apparatus, device and storage medium for generating entity relationship data

By identifying key-value blocks and their values in web page source code, the method addresses limitations of existing entity relationship data extraction, improving data volume and reducing human effort across diverse web pages.

CN109325201BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201810928930.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-08-15
Publication Date
2025-07-15
Estimated Expiration
2038-08-15

AI Technical Summary

Technical Problem

In the extraction of entity relationship data, the prior art has problems such as small amount of data, poor timeliness, high labor costs and poor universality.

Method used

By obtaining web page source code data, identifying key-value blocks and subject values, and generating entity relationship data, the automated method does not require human maintenance and is suitable for various web page structures.

Benefits of technology

It improves the output of entity relationship data and web page universality, reduces labor costs, and can extract large-scale entity relationship data from massive Internet web pages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109325201B_ABST
    Figure CN109325201B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, apparatus, device, and storage medium for generating entity relationship data. The method includes: obtaining web page source code data corresponding to a target web page; identifying at least one key-value block in the web page source code data, where the key-value block includes at least one key-value pair; identifying a main value corresponding to the at least one key-value block in the web page source code data; and generating entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block. Through the technical solution of the present invention, the web page generality can be improved, the labor cost can be reduced, and the output of entity relationship data can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to information processing technologies, and in particular, to a method, apparatus, device, and storage medium for generating entity relationship data. Background Art

[0002] Entity relationship data, also known as SPO triple data, refers to a triple composed of an entity pair (subject S - object O pair) and the relationship (P) between them. Entity relationships are a key component of knowledge graphs. From the perspective of knowledge graph construction, entity relationship mining can enrich the relationship knowledge in the graph and construct the association relationships between entities; from the perspective of product applications, entity relationships can directly meet users' search needs for knowledge - related content. For example, when searching for "the wife of a certain star", the entity relationship data can directly give the answer. On the other hand, it can also recommend related knowledge to users based on entity relationships, providing users with an information - extended reading experience. For example, when searching for the name of a certain celebrity, related other entities of this celebrity can be recommended to users through entity relationships.

[0003] In the prior art, entity relationship mining is mainly carried out through the following two methods:

[0004] Among them, the first method is to extract from encyclopedic websites. According to the characteristics of encyclopedic websites having a good structure and very standardized data, entity relationships are directly extracted from the information boxes or property tables (a web page structure used to describe entity attributes under an entity in an encyclopedic website) of encyclopedic websites. Utilizing the characteristics of the simple and stable structure of encyclopedic websites, several typical pages are sampled and labeled from the encyclopedic websites to be extracted. Then, one or more patterns represented in the form of class xpath are automatically constructed for these pages through a pattern learning algorithm, and then applied to other detailed pages of the website to achieve extraction.

[0005] The second method is the extraction method of generating wrappers (templates) for websites. By analyzing information such as the structure and HTML tags of the website to be extracted, corresponding wrappers are constructed, and these wrappers are used to extract entity relationships from the web page. For general regular pages, wrappers usually rely on manual writing of xpath and CSS selector expressions using regular expressions to extract elements in the web page.

[0006] The defects of the prior art are as follows: The first method can extract a small amount of data, and the data timeliness is not strong; the second method has a high labor cost and poor generality. Summary of the Invention

[0007] An embodiment of the present invention provides a method, apparatus, device, and storage medium for generating entity relationship data, so as to improve the universality of web pages, reduce labor costs, and increase the output of entity relationship data.

[0008] In a first aspect, an embodiment of the present invention provides a method for generating entity relationship data, including:

[0009] Obtaining web page source code data corresponding to a target web page;

[0010] Identifying at least one key-value block in the web page source code data, where the key-value block includes at least one key-value pair;

[0011] Identifying a main value corresponding to the at least one key-value block in the web page source code data;

[0012] Generating entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block.

[0013] In a second aspect, an embodiment of the present invention further provides an apparatus for generating entity relationship data, including:

[0014] A source code acquisition module for obtaining web page source code data corresponding to a target web page;

[0015] A key-value block identification module for identifying at least one key-value block in the web page source code data, where the key-value block includes at least one key-value pair;

[0016] A main value identification module for identifying a main value corresponding to the at least one key-value block in the web page source code data;

[0017] A data generation module for generating entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block.

[0018] In a third aspect, an embodiment of the present invention further provides a computer device, including:

[0019] One or more processors;

[0020] A storage device for storing one or more programs;

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating entity relationship data as described in any one of the embodiments of the present invention.

[0022] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating entity relationship data as described in any one of the embodiments of the present invention.

[0023] In the embodiment of the present invention, by obtaining the web page source code data corresponding to the target web page, and identifying at least one key-value block included in the web page source code data and the corresponding main body value of each key-value block, entity relationship data corresponding to the target web page is generated according to each key-value block and its corresponding main body value. Since entity relationships are identified from the web page source code data, there are no restrictions on web page types, web page structures, websites, etc., and entity relationship data can be automatically extracted from web pages without excessive manual maintenance. At the same time, a large amount of entity relationship data can be obtained by extracting from a large number of Internet web pages, thereby improving the universality of web pages, reducing labor costs, and increasing the output of entity relationship data. Description of the Drawings

[0024] Figure 1a is a flowchart of a method for generating entity relationship data provided in Embodiment 1 of the present invention;

[0025] Figure 1b is a flowchart of a method for preprocessing web page data applicable to Embodiment 1 of the present invention;

[0026] Figure 1c is a flowchart of a method for identifying main body values based on queries applicable to Embodiment 1 of the present invention;

[0027] Figure 1d is a flowchart of a method for identifying main body values applicable to Embodiment 1 of the present invention;

[0028] Figure 2a is a flowchart of a method for generating entity relationship data provided in Embodiment 2 of the present invention;

[0029] Figure 2b is a flowchart of a method for identifying key-value blocks applicable to Embodiment 2 of the present invention;

[0030] Figure 2c is a structural diagram of a semi-structured SPO data extraction system applicable to Embodiment 2 of the present invention;

[0031] Figure 3 is a structural diagram of a device for generating entity relationship data provided in Embodiment 3 of the present invention;

[0032] Figure 4 is a structural diagram of a computer device provided in Embodiment 4 of the present invention. Detailed Embodiments

[0033] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that, for the sake of convenience of description, only the parts related to the present invention rather than all the structures are shown in the drawings.

[0034] Additionally, it should be noted that, for the sake of convenience of description, only the parts related to the present invention rather than all the content are shown in the drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0035] Embodiment 1

[0036] Figure 1a As shown in the flowchart of a method for generating entity relationship data provided in Embodiment 1 of the present invention, this embodiment is applicable to the situation of obtaining entity relationship data from a web page. This method can be executed by the entity relationship data generation device provided in the embodiments of the present invention. The device can be composed of hardware and / or software and is generally integrated in a computer device. As shown in Figure 1a, the method of this embodiment specifically includes:

[0037] S110. Obtain the web page source code data corresponding to the target web page.

[0038] In this embodiment, the target web page can be any web page on the Internet, and the web page source code data can be the source code data of the target web page. Since there are all kinds of web pages on the Internet, the obtained web page can be preprocessed first to improve the effectiveness and reliability of the obtained web page source code data.

[0039] Specifically, obtaining the web page source code data corresponding to the target web page may include: obtaining the source code data corresponding to the target web page in the web page library according to the Uniform Resource Locator (URL) of the target web page as the source code data to be verified; if it is determined that the source code data to be verified does not meet the web page filtering conditions, then using the source code data to be verified as the web page source code data of the target web page. Among them, the source code data corresponding to multiple web pages are pre-stored in the web page library; the web page filtering conditions include but are not limited to filtering conditions such as the website where the web page is located being a blacklist website, the quality rating of the web page being less than a preset threshold, the language of the web page being a foreign language, the web page being a pornographic web page, and the type of the web page being a picture type, etc.; the source code data can be called ulpack data, and this data can be obtained from the web page library through tools such as wdbtools. In addition, the filtering of the source code data to be verified can be implemented using tools such as Nlpc antiporn.

[0040] To obtain the entity relationship data contained in the web page more comprehensively, optionally, after obtaining the web page source code data corresponding to the target web page, it further includes: obtaining at least one query formula corresponding to the Uniform Resource Locator of the target web page in the click display log of the search engine, and associating the obtained at least one query formula with the web page source code data.

[0041] Among them, the click display log can be the log recorded for the web page that the user clicks and opens among the displayed web pages after inputting the query formula. The query formula can be query data. The purpose of obtaining at least one query data corresponding to the URL of the target web page is to provide more references for the identification of the subject value in the subsequent steps. In addition, the obtained query formula can also be identified for violating query formulas, so as to further filter out violating web pages. Finally, by merging the query formulas corresponding to the same URL and their corresponding web page source code data, the association relationship between the query formulas and the web page source code data is established.

[0042] Take a practical example, such as Figure 1b shown in the process schematic diagram, the preprocessing of the URL of the obtained target web page can be completed through three steps, including:

[0043] Obtaining ulpack data: Obtaining the ulpack data corresponding to the URL of the target web page in the web page library;

[0044] Web page filtering: After obtaining the data in the web page library, filtering according to information such as the website where the web page is located, the quality of the web page, the language of the web page, whether it is a pornographic web page, and the type of the web page;

[0045] Filter pornographic queries and merge queries: Obtain the search queries of the URL, identify pornographic queries based on the queries of the URL, further filter out pornographic web pages, and merge the queries and web page data after filtering out pornographic queries.

[0046] S120. In the web page source code data, identify at least one key-value block, where the key-value block includes at least one key-value pair.

[0047] In this embodiment, the key-value pair can be KV type text data, and at least one KV type text data can form a key-value block. For example, if the web page information includes: "Gender: Female", it can be parsed from the web page source code data and then identified as a key-value pair. A key-value pair can include a key name and a key value. Among them, the key name (K) can be, for example, "Gender", and the key value (V) can be, for example, "Female".

[0048] Specifically, in the web page source code data, the key-value pair can be used as the minimum recognition unit to identify at least one key-value pair, and then these key values are divided according to the preset division rules to obtain at least one key-value block. The advantage of this setting is that it can reduce the processing objects in the subsequent steps and improve the processing efficiency.

[0049] Among them, the methods for identifying KV type text data from the web page source code data include, but are not limited to, tool identification, identification based on existing recognition results, and identification through specific tags, etc.

[0050] S130. In the web page source code data, identify the main values corresponding to at least one key-value block.

[0051] In this embodiment, the main value corresponding to the key-value block can be the main body (S) in the entity relationship data. The identification of the main value can be to find a suitable web page node as the main value for the key-value block according to the web page type, the type of the key-value block, and the xpath template of the target web page site. Among them, xpath is the path language of the subset (Extensible Markup Language, XML) of the Standard Generalized Markup Language, and it is a language used to determine the position of a certain part in the XML document.

[0052] Among them, the methods for separately identifying the main values corresponding to each key-value block include, but are not limited to, main value identification based on the entity page, main value identification based on strong style nodes, main value identification based on the whitelist, main value identification based on the query type, and main value identification based on the site template, etc.

[0053] For example, if the web page information includes: "Gender: Female, Age: 18 years old, Graduation Date: June 2018", and the corresponding person's name is "Zhang Xiaoting", then the first set of detailed information can be used as a key-value block, and "Zhang Xiaoting" can be identified as the main value corresponding to this key-value block.

[0054] Optionally, in the web page source code data, identifying the main value corresponding to at least one key-value block may include: if it is determined that the currently processed target key-value block is the main key-value block, and the web page source code data includes entity page nodes that meet the first label condition, then according to the entity page scoring rule, determine whether the target web page is an entity page; if so, use the text data corresponding to the entity page node as the main value of the target key-value block; where the main key-value block is the key-value block with the largest number of key-value pairs among at least one key-value block corresponding to the web page source code data.

[0055] Exemplarily, the main value identification based on the entity page includes two steps. First, determine whether the web page is an entity page, and then identify the main value of the entity page. Among them, the method for determining the entity page can be to perform weighted scoring according to the text features in the page, and if the score exceeds the threshold, it is considered an entity page type web page.

[0056] In a specific example, the text features in the page include whether there are certain keywords in the page (such as: introduction, overview, etc.), whether there is text describing scores (such as: scoring, score, etc.), the title length, whether the title coincides with the text in the page or the word frequency, etc. The weighted scoring method is to assign different weights to each text feature (typically, the weights can be set manually by sampling and statistically analyzing the discrimination of each feature) for summation scoring.

[0057] After determining that the target web page is an entity page, the specific identification of the main value of the entity page can be to directly use the text data corresponding to the entity page node in the web page source code data as the main value of the currently identified target key-value block. Among them, the entity page node can be, for example, in the web page source code data with <h1>Nodes of the label. Additionally, it should be emphasized that the target key-value block currently being recognized should be the key-value block with the largest number of key-value pairs among the multiple key-value blocks corresponding to the web page source code data, that is, the primary key-value block. The purpose of this setting is to avoid the content information corresponding to some small key-value blocks from being misrecognized as the content corresponding to the main title of the target web page, thereby improving the recognition accuracy of the main value.

[0058] Optionally, in the web page source code data, recognizing the main value corresponding to at least one key-value block may include: according to the page position of the target key-value block being processed in the target web page, searching forward in the web page source code data for strong style nodes that meet the second label condition; if a strong style node is found and the xpath of the strong style node is inconsistent with the xpath corresponding to the target key-value block, then the text data corresponding to the strong style node is used as the main value of the target key-value block.

[0059] Exemplarily, in the recognition of the main value based on strong style nodes, mainly according to the position of the key-value block, search forward in the web page for strong style nodes that meet the second label condition, and use the text data corresponding to the strong style nodes that meet certain rules as the main value of the key-value block. Among them, the strong style node can be, for example, in the web page source code data with <strong>Label or <h1>~ <h7>Nodes of the label. If the XPath of the strong style node is inconsistent with the XPath corresponding to this key-value block, it indicates that the strong style node has a higher level than this key-value block. Since this strong style node starts from the position in the web page where this key-value block is located and is searched forward, it is very likely to be the title information of the content corresponding to this key-value block. Therefore, the text data corresponding to this strong style node is used as the main value of this key-value block, reducing the misrecognition rate of the main value.

[0060] Optionally, in the web page source code data, identifying the main value corresponding to at least one key-value block may include: matching the key name of the key-value pair included in the currently processed target key-value block with a set whitelist; if it is determined that the target key name included in the target key-value block matches the whitelist, obtaining the target key value corresponding to the target key name as the main value of the target key-value block.

[0061] Exemplarily, since the key names of some key-value pairs in the key-value block have obvious features, for example, the object values corresponding to the entity relationships of the name type have obvious features, the key values corresponding to the key names with obvious features can be directly used as the main value. For example, key names with obvious features such as "name" and "movie name" can be preset in the whitelist. If the key-value block includes the key-value pair "name: Zhang Xiaoting", the key value "Zhang Xiaoting" corresponding to the key name "name" can be directly used as the main value of this key-value block.

[0062] Optionally, in the web page source code data, identifying the main value corresponding to at least one key-value block may include: if it is determined that the currently processed target key-value block is the main key-value block, determining the target token according to the word frequency of each token in at least one query formula associated with the web page source code data; searching for at least one query formula node in the web page source code data whose text data and the target token meet the similarity condition; if the XPath of the found query formula node is different from the XPath corresponding to the currently processed target key-value block, and in the target web page, the page position of the query formula node and the page position of the target key-value block meet the set distance condition, then using the text data of the query formula node as the main value of the target key-value block.

[0063] Exemplarily, the query-based subject value recognition is a subject value recognition strategy designed for the primary key value block. Since the query of a URL should have a relatively high relevance to the content of this page, and at the same time, the primary key value block of a web page also contains the main content information of the page, the content described by the information of the query and the primary key value block is very likely to be relatively similar. Therefore, the subject value of the primary key value block can be recognized according to the information contained in the query, that is, the target word segmentation. The input of the query-based subject value recognition is the word segmentation result of the query and the corresponding word frequency. Through the above recognition strategy, if there are nodes that meet the conditions, the recognized query nodes are given, and the text data corresponding to the query nodes is used as the subject value of the primary key value block, otherwise the recognition fails.

[0064] Among them, the purpose of setting the page position of the query node to meet the set distance condition with the page position of the target key value block is to ensure that the query node is not too far from the target key value block, reducing the probability of misrecognition.

[0065] Optionally, determining the target word segmentation according to the word frequency of each word segmentation in at least one query associated with the web page source code data may include: performing word segmentation processing on at least one query and calculating the word frequency of each word segmentation; if it is determined according to the word frequency calculation result that at least two word segmentations meet the word splicing condition, splicing at least two word segmentations to generate a new word segmentation and updating the word frequency corresponding to the new word segmentation; if the word frequency difference between the first-ranked word segmentation and the second-ranked word segmentation determined according to the sorting result of the word frequency of each word segmentation meets the word frequency threshold condition, taking the first-ranked word segmentation as the target word segmentation.

[0066] Exemplarily, a word segmentation tool can be used to perform word segmentation processing on each query, and then calculate the word frequency of each word segmentation according to the number of times each word segmentation appears. If the word frequencies of two of the word segmentations are both lower than the preset minimum threshold, the two word segmentations are spliced to form a new word segmentation, and the word frequency corresponding to each word segmentation is recalculated. Finally, according to the word frequency result, it is determined whether the word frequency difference between the word segmentation with the highest word frequency and the word segmentation with the second highest word frequency is greater than the preset threshold, that is, meets the word frequency threshold condition. If so, it means that the word segmentation with the highest word frequency is the word segmentation that can represent the main content of the query, and this word segmentation is taken as the target word segmentation to be compared with the content in the primary key value block, so as to finally determine the subject value corresponding to the primary key value block.

[0067] For the sake of easy understanding, the query-based subject value recognition process can be represented by a specific flowchart, for example Figure 1c As shown in the figure, specifically, first perform word segmentation on the query formula and calculate the corresponding word frequencies, then perform word splicing and recalculate the corresponding word frequencies. Determine whether the word frequencies corresponding to the two word segments with the highest word frequencies are similar. If so, the subject value cannot be identified through this query formula, so there is no output; if not, obtain the nodes in the web page that are similar to the word segment with the highest word frequency as candidate nodes. Determine whether there are such candidate nodes. If not, there is no output and this method fails. If so, after sorting each candidate node, sequentially determine whether each candidate node has the same xpath as the target key-value block being processed currently, and whether the distance between the position of the candidate node in the web page interface and the position of the target key-value block in the web page interface is too far. If one of them is yes, delete the candidate node. If both are no, determine the candidate node as the node where the subject value of the target key-value block is located, and the text data corresponding to this node is the subject value of the target key-value block.

[0068] Optionally, in the web page source code data, identifying the subject value corresponding to at least one key-value block may include: determining the target site corresponding to the target web page according to the uniform resource locator of the target web page; obtaining at least one candidate template stored in advance corresponding to the target site, and identifying the subject value corresponding to the key-value block being processed through the candidate template; wherein, the candidate templates in the target site are generated through the recognition results after performing key-value pair recognition on multiple web pages of the target site.

[0069] Exemplarily, the identification of the subject value based on the site template is to identify the subject value corresponding to the key-value block according to the site template data. It mainly includes two steps: obtaining candidate templates from the site template; using the candidate templates to identify the subject value. Wherein, the target site may be the site to which the target web page belongs. Since the URL of each web page carries the information of the site where the web page is located, therefore, the target site corresponding to the target web page can be identified according to this URL. Specifically, some sites may store the corresponding recognition templates in advance, so the subject value corresponding to the key-value block can be directly obtained by calling this template.

[0070] S140. Generate entity relationship data corresponding to the target web page according to the key-value block and the subject value corresponding to the key-value block.

[0071] The key-value blocks and their subject values in this embodiment may include multiple S-KV data, where S is the subject value of the key-value block, K is the key name of the key-value pair in the key-value block, and V is the key value of the key-value pair in the key-value block. Since there is a corresponding relationship between the S-KV data and the entity relationship data, therefore, the key-value blocks in the target web page can be split and recombined with the corresponding subject values to generate entity relationship data corresponding to the target web page.

[0072] Optionally, entity relationship data corresponding to the target web page can be generated according to the key-value block and the main body value corresponding to the key-value block, which may include: combining each key-value pair included in the key-value block with the main body value corresponding to the key-value block to construct triple data; using the key name included in the triple data as the main-object relationship value and the key value corresponding to the key name as the object value to generate entity relationship data.

[0073] Exemplarily, after combining a key-value pair and the main body value corresponding to its key-value block, an SPO triple data can be constructed. Among them, the main body value corresponding to the key-value block to which the key-value pair belongs is used as S in the SPO triple data, that is, the main body value in the entity relationship data; the key name of the key-value pair is used as P in the SPO triple data, that is, the entity relationship value in the entity relationship data; the key value of the key-value pair is used as O in the SPO triple data, that is, the object value in the entity relationship data, so as to finally generate entity relationship data.

[0074] The technical solution of this embodiment obtains the web page source code data corresponding to the target web page, identifies at least one key-value block and the main body value corresponding to each key-value block included in the web page source code data, and generates entity relationship data corresponding to the target web page according to each key-value block and its corresponding main body value. Since the entity relationship is identified from the web page source code data, it is not restricted by the web page type, web page structure, site, etc., and can automatically extract entity relationship data from the web page without excessive manual maintenance. At the same time, a large amount of entity relationship data can be obtained by extracting from a large number of Internet web pages, thereby improving the web page generality, reducing the labor cost, and increasing the output of entity relationship data.

[0075] Based on the above embodiments, optionally, one or more of the above five main body value identification methods for the target key-value block can be selected for main body value identification. Among them, if multiple methods are used for identification, the multiple methods can be combined in a certain order for comprehensive identification to improve the identification accuracy and success rate of the main body value. For example, after combining these five methods, the following methods can be used Figure 1d For the process shown, identify the main value of the key-value block. The specific identification order is: identification of the main value based on the main value of the entity page, identification of the main value based on strong style nodes, identification of the main value based on the whitelist, identification of the main value based on queries, and identification of the main value based on the site template. For the entity page, only the first identification method can be used to identify the main value. If the main value is identified, the corresponding result is output. If the main value is not identified, there is no output directly, that is, it is determined that the main value of the key-value block has not been successfully identified. For the identification of the main value in non-entity pages, in addition to the first identification method, as long as one of the identification methods identifies the main value, that is, there is an identification result, the result is output. Otherwise, continue to use the next identification method for identification until it is finally determined that the main value has indeed not been identified, and then it is determined that the main value of the key-value block has not been successfully identified.

[0076] Embodiment 2

[0077] Figure 2a The flowchart shown in this figure is a schematic flowchart of a method for generating entity relationship data provided in Embodiment 2 of the present invention. This embodiment is specific based on the above embodiment. In this embodiment, in the web page source code data, identifying at least one key-value block is further optimized to include: using a basic parsing tool to parse the web page source code data to obtain at least one basic key-value pair and adding it to the key-value pair set; performing key-value pair expansion on the basic key-value pair to obtain at least one expanded key-value pair and adding it to the key-value pair set; performing a merging process on the key-value pairs included in the key-value pair set to obtain at least one key-value block.

[0078] Correspondingly, the method of this embodiment includes:

[0079] S210. Obtain the web page source code data corresponding to the target web page.

[0080] S220. Use a basic parsing tool to parse the web page source code data to obtain at least one basic key-value pair and add it to the key-value pair set.

[0081] Among them, the basic parsing tool can be a KV parsing tool. The purpose of parsing the web page source code data corresponding to the target web page is to obtain at least one key-value pair included therein as the basic key-value pair and place it in a preset key-value pair set.

[0082] Exemplarily, since some parsing tools can automatically identify some data conforming to the KV type in the web page source code data, therefore, the basic parsing tool can be first used to roughly identify the key-value pairs included in the target web page.

[0083] For an actual example, the mars parsing tool can be used to parse ulpack data, and extract the text data described in the KV form within the web page from the parsing results, or the simple

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] <h1>Label <strong>Labels, < / strong> < / h1> The text data corresponding to the labels are used as basic key-value pairs. S230. Expand the basic key-value pairs to obtain at least one expanded key-value pair and add it to the key-value pair set. In this embodiment, due to certain limitations of the basic parsing tool, the number of key-value pairs that can be obtained from the data parsing result is limited. Therefore, it is also necessary to further identify other key-value pairs included in the target web page based on the obtained basic key-value pairs. Among them, the purpose of key-value pair expansion is to retrieve the content that cannot be parsed by the basic parsing tool. The ways of its key-value pair expansion include, but are not limited to, expanding the KV type text with the same xpath as the parsing result, and expanding the KV type text in specific HTML tags, etc. Optionally, expanding the basic key-value pairs to obtain at least one expanded key-value pair and adding it to the key-value pair set may include: In the web page source code data, obtain the basic xpath of the basic node that matches the basic key-value pair, and search for the expanded node with the same xpath as the basic xpath; Obtain the text data corresponding to the expanded node as the expanded key-value pair; And / or, in the web page source code data, obtain the basic html tag of the basic node that matches the basic key-value pair; Determine at least one expanded html tag according to the basic html tag, and search for the expanded node that matches the expanded html tag in the web page source code data; Obtain the text data corresponding to the expanded node as the expanded key-value pair. In the first method, since the probability of identifying other key-value pairs under the same xpath as the node where the already identified basic key-value pair is located is relatively high, therefore, the text data corresponding to other nodes under the same xpath as the node where the basic key-value pair is located can be obtained to expand the key-value pair and obtain more key-value pairs. In the second method, since each web page source code data contains multiple html tags, such as <strong>Labels, etc., the HTML tags corresponding to the recognized basic key-value pairs. There is a relatively high probability that other key-value pairs exist under this type of tag. Therefore, the key-value pairs can also be expanded by obtaining the HTML tags corresponding to the basic key-value pairs, that is, the basic HTML tags, and the text data corresponding to all matching nodes under the same type of tags.

[0090] The above two methods can be used separately or in combination, and there is no limitation here.

[0091] In an alternative implementation of this embodiment, optionally, after expanding the basic key-value pairs to obtain at least one expanded key-value pair and adding it to the key-value pair set, it may further include: performing a deduplication process on the key-value pairs included in the key-value pair set. Specifically, since the expanded key-value pairs may be repeated with the key-value pairs recognized by the tool, or there may be conflicts in node information, therefore, it is necessary to perform an integration and deduplication process on the expanded result.

[0092] In a specific example, the conflict of node information may include conflicts in location information (the location of the node in the web page) or text information, etc. For example, when the tool recognizes key-value pairs, there may be a location deviation, resulting in a conflict in the location information corresponding to the same key-value pair; or, for example, among two or more recognized key-value pairs, the same key name corresponds to different key values, resulting in a conflict in text information.

[0093] S240. Perform a merging process on the key-value pairs included in the key-value pair set to obtain at least one key-value block, where the key-value block includes at least one key-value pair.

[0094] In order to reduce the trouble of separately identifying the main value for each key-value pair in the subsequent steps and improve the generation efficiency of entity relationship data, in this embodiment, all the recognized key-value pairs are merged according to a preset merging rule, so that only the key-value blocks need to be identified as a unit in the subsequent main value identification, and only one main value needs to be identified for each key-value block. Specifically, the same number of key-value pairs or different numbers of key-value pairs can be allocated to each key-value block, and there is no limitation here.

[0095] Optionally, performing a merging process on the key-value pairs included in the key-value pair set to obtain at least one key-value block may include: locating the page positions of the key-value pairs in the key-value pair set in the target web page; merging at least two key-value pairs with continuous page positions into the same key-value block.

[0096] Exemplarily, multiple key-value pairs with consecutive page positions of the key-value pairs in the target web page, that is, there is no other text between the page positions where the key-value pairs are located, can be merged into one key-value block. Each web page can be merged into one or more key-value blocks, where the key-value block containing the most key-value pairs is the main key-value block.

[0097] Optionally, after merging the key-value pairs included in the key-value pair set to obtain at least one key-value block, it may further include: filtering the key-value pairs included in at least one key-value block according to the key-value pair filtering rule; filtering at least one key-value block according to the key-value block filtering rule.

[0098] Exemplarily, key-value filtering includes two granularities, namely the key-value pair granularity and the key-value block granularity. Specifically, key-value pair granularity filtering can be performed first, and then key-value block granularity filtering. The beneficial effect of filtering is that some invalid key-value pairs or key-value blocks can be deleted, thereby improving the effectiveness and reliability of the entity relationship data.

[0099] Among them, for the filtering of the key-value pair granularity, it is mainly filtered according to the length of the key-value pair, the text type, and the symbols therein. For example, a key-value pair with a too long key name or a too long key value in the key-value pair can be regarded as an invalid key-value pair and thus needs to be filtered out. Another example is that if the text corresponding to the key-value pair contains unrecognizable symbols or is all symbols, it should also be regarded as an invalid key-value pair.

[0100] In addition, for the filtering of the key-value block granularity, it is mainly to filter out the key-value blocks without entity relationship significance in the key-value block. For example, key-value blocks without practical significance such as "Warm reminder: ×××", and another example is text such as "Skill A: ×××" which is a glossary.

[0101] In a practical example, the entire processing process of key-value block recognition can be carried out according to the Figure 2b flow chart shown. First, use the mars parsing tool to parse the web page content, and perform expansion, deduplication, splitting, and filtering of KV according to the parsed KV data and web page information, and output the effective KV data in the final web page in units of KV blocks, where KV block splitting is the merging of key-value pairs.

[0102] S250. In the web page source code data, identify the main values corresponding to at least one key-value block.

[0103] S260. Generate entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block.

[0104] The technical solution of this embodiment obtains at least one basic key-value pair by parsing the web page source code data corresponding to the obtained target web page, expands the basic key-value pair obtained by using the basic parsing tool, and then adds both the basic key-value pair and the expanded key-value pair to the key-value pair combination. The key-value pairs included in the key-value pair set are merged to obtain at least one key-value block. Finally, in combination with the main values corresponding to the at least one key-value block identified in the web page source code data, entity relationship data corresponding to the target web page is generated. By using the basic key-value pairs obtained by parsing the web page source code data and the expanded key-value pairs expanded based on the basic key-value pairs, more key-value pairs can be obtained for each web page, thereby realizing the acquisition of a large amount of entity relationship data from a vast number of Internet web pages, which can not only meet the retrieval and recommendation requirements of popular entities, but also well solve the coverage problem of long-tail entities.

[0105] Based on the above embodiments, in order to improve the effectiveness of the finally generated entity relationship data, the obtained key-value blocks and their corresponding main values can be filtered. Optionally, after identifying the main values corresponding to at least one key-value block in the web page source code data, it may further include: using at least one statistical verification template and / or at least one rule verification template to filter at least one key-value block according to the main values corresponding to the at least one key-value block.

[0106] Exemplarily, filtering the key-value blocks is to perform quality control on the key-value blocks generated upstream and their corresponding main values, mainly including two methods: statistical-based verification and rule-based verification, filtering out the data that does not meet the conditions, so as to output high-quality key-value blocks.

[0107] Specifically, the statistical verification template is mainly used to filter errors that are difficult to solve under the information of a single page, and can perform statistical filtering from two granularities of key-value pairs and key-value blocks. In addition, the rule verification template is mainly based on the rules set by humans to process the S-KV data, that is, the key-value pairs (or key-value blocks) and their corresponding main values, and filter S and KV respectively.

[0108] Optionally, it may further include: obtaining, in each web page in the target site corresponding to the uniform resource locator of the target web page, the recognized main body value corresponding to the key-value block as the processed main body value; if the number of target processed main body values with the same xpath exceeds the first quantity threshold, and constructing at least one alternative statistical verification template according to the xpath of the target main body value and the key-value blocks corresponding to the target main body values respectively; obtaining the key-value blocks corresponding to the alternative statistical verification templates respectively; if the number of identical key-value blocks among the multiple key-value blocks corresponding to a target alternative statistical verification template exceeds the second quantity threshold, deleting the target alternative statistical verification template in the alternative statistical verification template to obtain the statistical verification template.

[0109] Exemplarily, the statistical verification template can be constructed through two verification steps: The first step is site-S xpath level statistical verification, that is, merging the site where the URL is located and the xpath of the main body value recognized by the URL as the key, establishing multiple initial templates based on the merging result, and matching the actual S-KV recognition result with each initial template, filtering out the initial templates that do not meet the conditions, that is, obtaining the alternative statistical verification templates; The second step is the S xpath-KV block level statistical verification template, that is, on the basis of the alternative statistical verification template, merging the S xpath, that is, the xpath of the main body value, and the xpath of the corresponding KV block, that is, the xpath of the key-value block, as the key, and deleting the templates with completely identical KV under multiple URLs after merging, that is, finally obtaining the statistical verification template. The finally obtained statistical verification template mainly filters the KV data included in the public information of the sidebar or bottom bar of the web page.

[0110] Based on the above embodiments, a semi-structured SPO data extraction system as shown in Figure 2c can be constructed. The system includes: a preprocessing unit 50, a KV recognition unit 60, an S recognition unit 70, and an S-KV quality control unit 80. Among them, the preprocessing unit 50 is used for preprocessing the input web page URL to obtain the web page source code data; the KV recognition unit 60 is used for recognizing the key-value blocks in the web page source code data; the S recognition unit 70 is used for recognizing the main body values in the web page source code data; the S-KV quality control unit 80 is used for filtering the S-KV data that does not meet the requirements, and then outputting the SPO triple data that meets the requirements.

[0111] The goal of this semi-structured SPO data extraction system is to implement a system that extracts and organizes information represented in the form of key-value pairs from web pages into data in the form of triples. That is, given the URL of a target web page as input, this system will identify the corresponding entities (S) for the entity relationships (P) and entity attribute values (O) identified by the target web page in the form of pairs, and then output them in the form of SPO triples, so as to realize the automatic extraction of SPO data, that is, entity relationship data, from any type of web page. As Figure 2c shown, after entering the URL through the semi-structured SPO data extraction system and using corresponding external data, such as click display logs required in the preprocessing process and site templates required in the process of identifying the main body values based on site templates, the corresponding SPO triple data is finally obtained.

[0112] Embodiment 3

[0113] Figure 3 is a schematic structural diagram of a device for generating entity relationship data provided in Embodiment 3 of the present invention. Refer to Figure 3 , the device for generating entity relationship data includes: a source code acquisition module 310, a key-value block identification module 320, a main body value identification module 330, and a data generation module 340. The following is a specific description of each module.

[0114] The source code acquisition module 310 is used to acquire the web page source code data corresponding to the target web page;

[0115] The key-value block identification module 320 is used to identify at least one key-value block in the web page source code data, where the key-value block includes at least one key-value pair;

[0116] The main body value identification module 330 is used to identify the main body value corresponding to the at least one key-value block in the web page source code data;

[0117] The data generation module 340 is used to generate entity relationship data corresponding to the target web page according to the key-value block and the main body value corresponding to the key-value block.

[0118] This embodiment provides a device for generating entity relationship data. By obtaining the web page source code data corresponding to the target web page, and identifying at least one key-value block and the corresponding body value included in the web page source code data, entity relationship data corresponding to the target web page is generated according to each key-value block and its corresponding body value. Since entity relationships are identified from the web page source code data, there are no restrictions on web page types, web page structures, websites, etc., and entity relationship data can be automatically extracted from web pages without excessive manual maintenance. At the same time, a large amount of entity relationship data can be obtained by extracting from a large number of Internet web pages, thereby improving the universality of web pages, reducing labor costs, and increasing the output of entity relationship data.

[0119] Optionally, the key-value block recognition module 320 may include:

[0120] The basic key-value acquisition sub-module is used to perform data parsing on the web page source code data by using a basic parsing tool, and obtain at least one basic key-value pair and add it to the key-value pair set;

[0121] The extended key-value acquisition sub-module is used to perform key-value pair extension on the basic key-value pair, and obtain at least one extended key-value pair and add it to the key-value pair set;

[0122] The key-value block determination sub-module is used to perform a merging process on the key-value pairs included in the key-value pair set to obtain the at least one key-value block.

[0123] Optionally, the extended key-value acquisition sub-module may specifically be used for:

[0124] In the web page source code data, obtain the basic xpath of the basic node that matches the basic key-value pair, and search for the extended node whose xpath is the same as the basic xpath; obtain the text data corresponding to the extended node as the extended key-value pair; and / or

[0125] In the web page source code data, obtain the basic html tag of the basic node that matches the basic key-value pair; determine at least one extended html tag according to the basic html tag, and search for the extended node that matches the extended html tag in the web page source code data; obtain the text data corresponding to the extended node as the extended key-value pair.

[0126] Optionally, the key-value block determination sub-module may specifically be used for:

[0127] Locate the page position of the key-value pair in the key-value pair set in the target web page;

[0128] Merge at least two key-value pairs with continuous page positions into the same key-value block.

[0129] Optionally, the key-value block recognition module 320 may further include:

[0130] A key-value pair filtering sub-module, configured to, after performing a merging process on the key-value pairs included in the key-value pair set to obtain the at least one key-value block, perform a filtering process on the key-value pairs included in the at least one key-value block according to a key-value pair filtering rule;

[0131] A key-value block filtering sub-module, configured to perform a filtering process on the at least one key-value block according to a key-value block filtering rule.

[0132] Optionally, the body value recognition module 330 may specifically be configured to:

[0133] If it is determined that the target key-value block being currently processed is the primary key-value block and the web page source code data includes entity page nodes that meet the first label condition, then determine whether the target web page is an entity page according to an entity page scoring rule;

[0134] If so, use the text data corresponding to the entity page node as the body value of the target key-value block;

[0135] Wherein, the primary key-value block is the key-value block with the largest number of key-value pairs among the at least one key-value block corresponding to the web page source code data.

[0136] Optionally, the body value recognition module 330 may specifically be configured to:

[0137] According to the page position of the target key-value block being currently processed in the target web page, search forward in the web page source code data for strong style nodes that meet the second label condition;

[0138] If the strong style node is found and the xpath of the strong style node is inconsistent with the xpath corresponding to the target key-value block, then use the text data corresponding to the strong style node as the body value of the target key-value block.

[0139] Optionally, the body value recognition module 330 may specifically be configured to:

[0140] Match the key names of the key-value pairs included in the target key-value block being currently processed with a set whitelist;

[0141] If it is determined that the target key name included in the target key-value block matches the whitelist, then obtain the target key value corresponding to the target key name as the body value of the target key-value block.

[0142] The generating device for the entity relationship data may further include:

[0143] A query-based acquisition module, configured to, after obtaining the web page source code data corresponding to a target web page, obtain at least one query corresponding to the uniform resource locator of the target web page in the click display log of a search engine, and associate the obtained at least one query with the web page source code data;

[0144] Correspondingly, the main value recognition module 330 may include:

[0145] A word segmentation determination sub-module, configured to, if it is determined that the target key-value block currently being processed is the main key-value block, determine the target word segmentation according to the word frequency of each word segmentation in the at least one query associated with the web page source code data;

[0146] A node search sub-module, configured to search for at least one query node in the web page source code data whose text data meets the similarity condition with the target word segmentation;

[0147] A main value determination sub-module, configured to, if the xpath of the found query node is different from the xpath corresponding to the target key-value block currently being processed, and in the target web page, the page position of the query node meets the set distance condition with the page position of the target key-value block, then use the text data of the query node as the main value of the target key-value block.

[0148] Optionally, the word segmentation determination sub-module may specifically be configured to:

[0149] Perform word segmentation processing on the at least one query, and calculate the word frequency of each word segmentation;

[0150] If it is determined according to the word frequency calculation result that at least two word segmentations meet the word splicing condition, splice the at least two word segmentations to generate a new word segmentation, and update the word frequency corresponding to the new word segmentation;

[0151] If the word frequency difference between the first-ranked word segmentation and the second-ranked word segmentation determined according to the sorting result of the word frequencies of each word segmentation meets the word frequency threshold condition, use the first-ranked word segmentation as the target word segmentation.

[0152] Optionally, the main value recognition module 330 may specifically be configured to:

[0153] Determine a target site corresponding to the target web page according to the uniform resource locator of the target web page;

[0154] Obtain at least one candidate template pre-stored corresponding to the target site, and identify the main value corresponding to the key-value block currently being processed through the candidate template;

[0155] Among them, the candidate templates in the target site are generated through the recognition results after identifying key-value pairs in multiple web pages of the target site.

[0156] Optionally, the apparatus for generating entity relationship data may further include:

[0157] A template filtering module, configured to, after identifying the main body values corresponding to the at least one key-value block in the web page source code data, use at least one statistical verification template and / or at least one rule verification template to filter the at least one key-value block according to the main body values corresponding to the at least one key-value block.

[0158] Optionally, the apparatus for generating entity relationship data may further include:

[0159] A main body value obtaining module, configured to respectively obtain, in each web page in the target site corresponding to the uniform resource locator of the target web page, the identified main body values corresponding to the key-value blocks as the processing main body values;

[0160] A template construction module, configured to, if the number of target processing main body values with the same xpath exceeds the first quantity threshold, and according to the xpath of the target main body values and the key-value blocks respectively corresponding to the target main body values, construct at least one alternative statistical verification template;

[0161] A key-value block obtaining module, configured to obtain the key-value blocks respectively corresponding to the alternative statistical verification templates;

[0162] A template determination module, configured to, if the number of identical key-value blocks among the multiple key-value blocks corresponding to a target alternative statistical verification template exceeds the second quantity threshold, delete the target alternative statistical verification template in the alternative statistical verification templates to obtain the statistical verification template.

[0163] Optionally, the data generation module 340 may specifically be configured to:

[0164] Respectively combine each key-value pair included in the key-value block with the main body value corresponding to the key-value block to construct triple data;

[0165] Use the key name included in the triple data as the main-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data.

[0166] The above product can execute the method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the executed method.

[0167] Embodiment 4

[0168] Figure 4 Schematic diagram of a computer device provided in Embodiment 4 of the present invention. Figure 4 The block diagram of an exemplary computer device 12 suitable for implementing the embodiments of the present invention is shown. Figure 4 The displayed computer device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0169] As Figure 4 shown, the computer device 12 is presented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0170] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0171] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0172] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used for reading and writing non-removable, non-volatile magnetic media ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing removable non-volatile disks (such as "floppy disks") and an optical disk drive for reading and writing removable non-volatile optical disks (such as CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data media interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0173] A program / utilities 40 having a set (at least one) of program modules 42 can be stored, for example, in a memory 28. Such program modules 42 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.

[0174] The computer device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the computer device 12, and / or communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the computer device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computer device 12 through a bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0175] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28. For example, it implements the method for generating entity relationship data provided in the embodiments of the present invention: obtaining web page source code data corresponding to a target web page; identifying at least one key-value block in the web page source code data, where the key-value block includes at least one key-value pair; identifying a main value corresponding to the at least one key-value block in the web page source code data; and generating entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block.

[0176] Embodiment Five

[0177] Example 5 of the embodiments of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for generating entity relationship data provided in all the inventive embodiments of the present application: obtaining web source code data corresponding to a target web page; identifying at least one key-value block in the web source code data, where the key-value block includes at least one key-value pair; identifying a main value corresponding to the at least one key-value block in the web source code data; and generating entity relationship data corresponding to the target web page according to the key-value block and the main value corresponding to the key-value block.

[0178] One or more computer-readable media can be used in any combination. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0179] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0180] The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including - but not limited to - wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0181] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0182] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.< / strong> < / h7> < / h1> < / strong> < / h1>

Claims

1. A method for generating entity relationship data, characterized in that Including: Obtain the web page source code data corresponding to the target web page; In the web page source code data, identify at least one key-value block, where the key-value block includes at least one key-value pair; In the web page source code data, identify the main body value corresponding to the at least one key-value block; Combine each key-value pair included in the key-value block with the main body value corresponding to the key-value block respectively to construct triple data; Use the key name included in the triple data as the main-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data, where the entity relationship data includes the main body value, the entity relationship value, and the object value; Among them, identifying at least one key-value block in the web page source code data includes: by parsing the web page source code data, identifying at least one key-value pair, and dividing the key-value pair according to a preset division rule to obtain at least one key-value block; The identifying the main body value corresponding to the at least one key-value block in the web page source code data includes: If it is determined that the target key-value block being currently processed is the main key-value block, and the web page source code data includes an entity page node that meets the first label condition, then judge whether the target web page is an entity page according to the entity page scoring rule; If so, use the text data corresponding to the entity page node as the main body value of the target key-value block; Among them, the main key-value block is the key-value block with the largest number of key-value pairs among at least one key-value block corresponding to the web page source code data.

2. The method according to claim 1, characterized in that, Identifying at least one key-value block in the web page source code data includes: Use a basic parsing tool to parse the web page source code data to obtain at least one basic key-value pair and add it to the key-value pair set; Perform key-value pair expansion on the basic key-value pair to obtain at least one extended key-value pair and add it to the key-value pair set; Perform a merging process on the key-value pairs included in the key-value pair set to obtain the at least one key-value block.

3. The method according to claim 2, characterized in that Performing key-value pair expansion on the basic key-value pair to obtain at least one extended key-value pair and add it to the key-value pair set includes: In the web page source code data, obtain the basic xpath of the basic node matching the basic key-value pair, and search for an extended node with the same xpath as the basic xpath; obtain the text data corresponding to the extended node as the extended key-value pair; and / or In the web page source code data, obtain the basic html tag of the basic node matching the basic key-value pair; according to the basic html tag, determine at least one extended html tag, and in the web page source code data, search for an extended node matching the extended html tag; obtain the text data corresponding to the extended node as the extended key-value pair.

4. The method according to claim 2, characterized in that Performing a merging process on the key-value pairs included in the key-value pair set to obtain the at least one key-value block includes: Locate the page positions of the key-value pairs in the key-value pair set in the target web page; Merge at least two key-value pairs with continuous page positions into the same key-value block.

5. The method according to claim 2, wherein After merging the key-value pairs included in the key-value pair set to obtain the at least one key-value block, the method further includes: Filtering the key-value pairs included in the at least one key-value block according to a key-value pair filtering rule; Filtering the at least one key-value block according to a key-value block filtering rule.

6. The method according to claim 1, characterized in that, After obtaining the web page source code data corresponding to the target web page, the method further includes: Obtaining at least one query formula corresponding to the uniform resource locator of the target web page in the click display log of the search engine, and associating the obtained at least one query formula with the web page source code data; Identifying a main value corresponding to the at least one key-value block in the web page source code data, including: If it is determined that the target key-value block being currently processed is a main key-value block, determining a target word segment according to the word frequency of each word segment in the at least one query formula associated with the web page source code data; Searching in the web page source code data for at least one query formula node whose text data meets the similarity condition with the target word segment; If the xpath of the found query formula node is different from the xpath corresponding to the target key-value block being currently processed, and in the target web page, the page position of the query formula node meets the set distance condition with the page position of the target key-value block, then using the text data of the query formula node as the main value of the target key-value block.

7. The method according to claim 6, wherein Determining a target word segment according to the word frequency of each word segment in the at least one query formula associated with the web page source code data, including: Performing word segment processing on the at least one query formula and calculating the word frequency of each word segment; If it is determined according to the word frequency calculation result that at least two word segments meet the word splicing condition, splicing the at least two word segments to generate a new word segment, and updating the word frequency corresponding to the new word segment; If the word frequency difference between the first-ranked word segment and the second-ranked word segment determined according to the sorting result of the word frequencies of each word segment meets the word frequency threshold condition, then using the first-ranked word segment as the target word segment.

8. The method according to any one of claims 1 to 7, characterized in that After identifying the main value corresponding to the at least one key-value block in the web page source code data, the method further includes: Using at least one statistical verification template and / or at least one rule verification template to filter the at least one key-value block according to the main value corresponding to the at least one key-value block.

9. The method according to claim 8, wherein The method further includes: Respectively obtaining, in each web page of the target site corresponding to the uniform resource locator of the target web page, the identified main value corresponding to the key-value block as the processing main value; If the number of target processing main values with the same xpath exceeds a first quantity threshold, and constructing at least one alternative statistical verification template according to the xpath of the target processing main value and the key-value blocks respectively corresponding to the target processing main value; Obtaining the key-value blocks respectively corresponding to the alternative statistical verification templates; If the number of identical key-value blocks among the multiple key-value blocks corresponding to a target alternative statistical verification template exceeds a second quantity threshold, then deleting the target alternative statistical verification template in the alternative statistical verification template to obtain the statistical verification template.

10. A method for generating entity relationship data, characterized in that Including: Obtain the web page source code data corresponding to the target web page; In the web page source code data, identify at least one key-value block, where the key-value block includes at least one key-value pair; In the web page source code data, identify the main body value corresponding to the at least one key-value block; Combine each key-value pair included in the key-value block with the main body value corresponding to the key-value block respectively to construct triple data; Use the key name included in the triple data as the main-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data, where the entity relationship data includes the main body value, the entity relationship value and the object value; Wherein, identifying at least one key-value block in the web page source code data includes: by parsing the web page source code data, identifying at least one key-value pair, and dividing the key-value pair according to a preset division rule to obtain at least one key-value block; The identifying the main body value corresponding to the at least one key-value block in the web page source code data includes: According to the page position of the target key-value block being processed in the target web page, forward search in the web page source code data for strong style nodes that meet the second label condition; If the strong style node is found and the xpath of the strong style node is inconsistent with the xpath corresponding to the target key-value block, then use the text data corresponding to the strong style node as the main body value of the target key-value block.

11. A method for generating entity relationship data, characterized in that, Includes: Obtain the web page source code data corresponding to the target web page; In the web page source code data, identify at least one key-value block, where the key-value block includes at least one key-value pair; In the web page source code data, identify the main body value corresponding to the at least one key-value block; Combine each key-value pair included in the key-value block with the main body value corresponding to the key-value block respectively to construct triple data; Use the key name included in the triple data as the main-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data, where the entity relationship data includes the main body value, the entity relationship value and the object value; Wherein, identifying at least one key-value block in the web page source code data includes: by parsing the web page source code data, identifying at least one key-value pair, and dividing the key-value pair according to a preset division rule to obtain at least one key-value block; The identifying the main body value corresponding to the at least one key-value block in the web page source code data includes: Match the key name of the key-value pair included in the target key-value block being processed with the set white list; If it is determined that the target key name included in the target key-value block matches the white list, then obtain the target key value corresponding to the target key name as the main body value of the target key-value block.

12. A method for generating entity relationship data, characterized in that, Includes: Obtain the web page source code data corresponding to the target web page; In the web page source code data, identify at least one key-value block, where the key-value block includes at least one key-value pair; In the web page source code data, identify the main body value corresponding to the at least one key-value block; Combine each key-value pair included in the key-value block with the corresponding body value of the key-value block to construct triple data; Use the key name included in the triple data as the subject-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data, where the entity relationship data includes a body value, an entity relationship value, and an object value; Among them, identifying at least one key-value block in the web page source code data includes: by parsing the web page source code data, identifying at least one key-value pair, and dividing the key-value pair according to a preset division rule to obtain at least one key-value block; Identifying the body value corresponding to the at least one key-value block in the web page source code data includes: Determine the target site corresponding to the target web page according to the uniform resource locator of the target web page; Obtain at least one candidate template pre-stored corresponding to the target site, and identify the body value corresponding to the currently processed key-value block through the candidate template; Among them, the candidate templates in the target site are generated through the recognition results after identifying key-value pairs for multiple web pages of the target site.

13. An apparatus for generating entity relationship data, characterized in that Include: A source code acquisition module for acquiring web page source code data corresponding to a target web page; A key-value block recognition module for parsing the web page source code data to identify at least one key-value pair, and dividing the key-value pair according to a preset division rule to obtain at least one key-value block, where the key-value block includes at least one key-value pair; A body value recognition module for identifying the body value corresponding to the at least one key-value block in the web page source code data; A data generation module for combining each key-value pair included in the key-value block with the corresponding body value of the key-value block to construct triple data; using the key name included in the triple data as the subject-object relationship value, and the key value corresponding to the key name as the object value to generate the entity relationship data, where the entity relationship data includes a body value, an entity relationship value, and an object value; The body value recognition module is specifically used for: If it is determined that the currently processed target key-value block is the main key-value block and the web page source code data includes entity page nodes that meet the first label condition, then judge whether the target web page is an entity page according to the entity page scoring rule; If so, use the text data corresponding to the entity page node as the body value of the target key-value block; Among them, the main key-value block is the key-value block with the largest number of key-value pairs among at least one key-value block corresponding to the web page source code data.

14. A computer device, characterized in that, The device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating entity relationship data as described in any one of claims 1-12.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method for generating entity relationship data as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Method for automatically collecting webpage content

    CN104933168A

  • A Deepdive-based field text knowledge extracting method

    CN107169079A

  • Open entity relationship extracting method based on sentence meaning structure model

    CN108363816A

  • Webpage data processing method and apparatus, query processing method and question-answering system

    CN104516949A