Information extraction method, device and equipment

By acquiring HTML data to construct a knowledge graph and using graph convolutional networks for reasoning, the problem of low efficiency in manual information extraction is solved, and efficient and accurate automated information extraction and real-time response are achieved.

CN120632241APending Publication Date: 2025-09-12CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510592367.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the existing technology, manual information extraction methods are time-consuming, difficult to adapt to the rapid growth of data volume, and have poor consistency and real-time performance of results, resulting in low efficiency and accuracy in information extraction from unstructured web page data.

Method used

By obtaining the HTML data of the target web page, using natural language processing to build a knowledge graph, and performing knowledge reasoning based on the graph convolutional network, the implicit relationship between entities is determined, and the knowledge graph is optimized to obtain information extraction results.

Benefits of technology

It realizes the automation and intelligence of information extraction, improves efficiency and accuracy, ensures the consistency of results, and supports real-time updates and responses to changes in web page data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632241A_ABST
    Figure CN120632241A_ABST
Patent Text Reader

Abstract

The invention discloses an information extraction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining hypertext markup language (HTML) data of a target webpage, and extracting unstructured data in the HTML data; natural language processing is carried out on the unstructured data, a knowledge graph is constructed, and nodes in the knowledge graph represent entities in the unstructured data; and performing knowledge reasoning on the knowledge graph based on the graph convolutional network, determining an implicit relationship between the entities, and optimizing the knowledge graph based on the implicit relationship to obtain an information extraction result of the target webpage. The HTML data of the target webpage are obtained and processed through an automatic process, the information extraction efficiency can be greatly improved, the challenge of mass data can be effectively handled, knowledge graph construction and reasoning are carried out based on natural language processing and the graph convolutional network, the standardization of the processing process is achieved, the consistency of results is ensured, and the processing efficiency is improved. The influence of subjective judgment is reduced, and the accuracy of information extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to an information extraction method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of information technology, the Internet has become one of the main channels for information dissemination. As the carrier of Internet information, web pages carry a massive amount of unstructured data in various forms, including text, images, and videos.

[0003] At present, we can understand the content of web pages through manual reading, and then manually extract information and build knowledge graphs to achieve information extraction of unstructured data on web pages, and effectively extract and utilize these unstructured data.

[0004] However, manual information extraction methods are not only time-consuming and difficult to adapt to the rapid growth of data volume, but also affect the processing results due to individual differences and subjective judgments, making it difficult to ensure the consistency of processing results and unable to achieve real-time updates and responses, resulting in low efficiency and accuracy in information extraction of unstructured data on web pages. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an information extraction method, apparatus, device and storage medium that can solve the current problem of low efficiency and accuracy in information extraction from unstructured data on web pages.

[0006] In a first aspect, an embodiment of the present application provides an information extraction method, the method comprising:

[0007] Obtaining Hypertext Markup Language (HTML) data of a target web page and extracting unstructured data from the HTML data;

[0008] Performing natural language processing on the unstructured data to construct a knowledge graph, wherein nodes in the knowledge graph represent entities in the unstructured data;

[0009] The knowledge graph is subjected to knowledge reasoning based on a graph convolutional network to determine the implicit relationships between the entities, and the knowledge graph is optimized based on the implicit relationships to obtain the information extraction result of the target web page.

[0010] Optionally, extracting unstructured data from the HTML data includes:

[0011] Parsing the HTML data and converting the HTML data into a Document Object Model (DOM) tree;

[0012] Traversing the DOM tree to perform structural analysis and determine the web page elements corresponding to the unstructured data in the HTML data;

[0013] The unstructured data is extracted from the web page elements.

[0014] Optionally, extracting the unstructured data from the webpage element includes:

[0015] In the case where the web page element includes a table, traversing the table in the web page element according to a preset splitting strategy to extract the table data;

[0016] A data storage unit corresponding to the table data is created, and the table data is stored in the data storage unit to obtain the unstructured data.

[0017] Optionally, performing natural language processing on the unstructured data to construct a knowledge graph includes:

[0018] Performing natural language processing on the unstructured data to extract entities in the unstructured data, semantic relationships between the entities, and attribute information of the entities;

[0019] A knowledge graph is constructed based on the entities, the relationships between the entities, and the attribute information of the entities, wherein the edges between the nodes in the knowledge graph represent the corresponding semantic relationships between the entities, and the attribute information is used to represent the attributes of the nodes or the edges.

[0020] Optionally, performing natural language processing on the unstructured data to extract entities in the unstructured data, semantic relationships between the entities, and attribute information of the entities includes:

[0021] Performing named entity recognition on the unstructured data, extracting entities from the unstructured data, and identifying attribute information related to the entities;

[0022] Semantic recognition is performed on the unstructured data to obtain semantic relationships between the entities.

[0023] Optionally, the performing of knowledge reasoning on the knowledge graph based on a graph convolutional network to determine implicit relationships between the entities, and optimizing the knowledge graph based on the implicit relationships to obtain information extraction results for the target web page includes:

[0024] Performing deduplication processing, disambiguation processing, and logical verification on the knowledge graph to obtain an optimized knowledge graph;

[0025] Performing knowledge reasoning on the optimized knowledge graph based on a graph convolutional network to mine implicit relationships between the entities;

[0026] The implicit relationship is added to the optimized knowledge graph to obtain the information extraction result of the target web page.

[0027] Optionally, after optimizing the knowledge graph based on the implicit relationship to obtain the information extraction result of the target webpage, the method further includes:

[0028] The information extraction result is stored in a database, and an index of the information extraction result is created in the database, where the index is used to query the information extraction result.

[0029] In a second aspect, an embodiment of the present application provides an information extraction device, applied to a controller, comprising:

[0030] An acquisition module, configured to acquire the Hypertext Markup Language (HTML) data of a target web page and extract unstructured data from the HTML data;

[0031] A construction module, configured to perform natural language processing on the unstructured data to construct a knowledge graph, wherein nodes in the knowledge graph represent entities in the unstructured data;

[0032] An optimization module is used to perform knowledge reasoning on the knowledge graph based on a graph convolutional network, determine the implicit relationship between the entities, and optimize the knowledge graph based on the implicit relationship to obtain the information extraction result of the target web page.

[0033] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0034] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0035] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0036] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.

[0037] From the above, it can be seen that this application obtains and processes the HTML data of the target web page through an automated process, which can greatly improve the efficiency of information extraction and effectively cope with the challenges of massive data. Moreover, the knowledge graph construction and reasoning based on natural language processing and graph convolutional networks not only realizes the standardization of the processing process, ensures the consistency of the results, and reduces the influence of subjective judgment, but also can deeply explore and determine the implicit relationships between entities, thereby significantly improving the accuracy and depth of information extraction. In addition, this automated and intelligent method makes it possible to update and respond to changes in web page data in real time, overcoming the real-time limitations of manual methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart showing an information extraction method according to an exemplary embodiment;

[0039] Figure 2 This is a schematic diagram of a parsing process of HTML data according to an exemplary embodiment;

[0040] Figure 3 is a schematic diagram showing a hierarchical structure of HTML data according to an exemplary embodiment;

[0041] Figure 4 is an example diagram showing a triple graph relationship according to an exemplary embodiment;

[0042] Figure 5 is a schematic diagram illustrating an example of an information extraction method according to an exemplary embodiment;

[0043] Figure 6 is a block diagram of an information extraction device according to an exemplary embodiment;

[0044] Figure 7 is a block diagram of an electronic device according to an exemplary embodiment;

[0045] Figure 8 The figure is a schematic diagram showing the hardware structure of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0046] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0047] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0048] First, let’s explain the terms mentioned in this application:

[0049] Java (Java programming language) is a widely used static object-oriented programming language specially designed for writing cross-platform applications. It has the characteristics of being powerful and easy to use.

[0050] HTML (Hypertext Markup Language) is a markup language that uses a series of specific tags to define the structure and content of web pages. This standardizes the previously disparate and independent documents on the web, giving them a consistent structure and presentation, and connecting scattered web resources into a logical whole. Specifically, HTML data is descriptive text composed of HTML commands. HTML commands precisely tell the browser how to interpret and display web page content, such as defining text styles, embedding graphics and animations, playing sounds, organizing table data, and creating links to other resources.

[0051] DOM (Document Object Model) is a programming interface for HTML and XML (eXtensible Markup Language) documents. When a browser receives HTML data, its first task is to parse the HTML data and construct a DOM tree. The DOM tree intuitively presents the internal logic of the HTML document in a hierarchical tree structure. In the DOM tree, each node is an independent object. Starting from the root node, child nodes are derived layer by layer until the end of each branch. These nodes together constitute a digital mapping of the web page content. With the rich interfaces and methods provided by DOM, developers can easily access and modify the structure, style, and content of nodes.

[0052] Natural Language Processing (NLP) is an interdisciplinary field dedicated to studying how humans and computers can effectively communicate using natural language (the language we use in everyday life). It integrates theories and methods from linguistics, computer science, and mathematics. The core goal of NLP is to develop computer systems, particularly the software systems, that can effectively implement natural language communication. With this goal in mind, NLP technology has been widely used in areas such as machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese optical character recognition (OCR).

[0053] Named Entity Recognition (NER), also known as "proper name recognition," refers to the identification of entities with specific meanings in text, such as names of people, places, organizations, and proper nouns. These entities often carry key information in the text. Therefore, NER technology has become an important foundational tool in many application fields, including information extraction, question-answering systems, syntactic analysis, and machine translation, and plays a crucial role in the practical application of natural language processing technology. To more clearly define the scope of recognition, NER tasks are often concretized into identifying three types of entities: entity types (such as the aforementioned names of people, places, and organizations), time types (such as specific time points and time periods), and numerical types (such as currency, percentages, and quantities).

[0054] The Bidirectional Long Short-Term Memory-Conditional Random Field (BLSTM-CRF) model combines two key components: a bidirectional LSTM (BiLSTM) and a conditional random field (CRF). It is a deep learning architecture that excels in sequence labeling tasks such as named entity recognition. By integrating a bidirectional LSTM network, the BiLSTM overcomes the limitation of the unidirectional LSTM, which only captures information from left to right (or front to back). It can process both forward and backward information in a text sequence, considering not only the context preceding a word but also the context following it. This allows for a more comprehensive and in-depth understanding of the text, laying the foundation for accurate identification of entity boundaries and types. The final output layer of the BiLSTM network is typically followed by a linear layer. This layer projects the hidden layer output generated by the BiLSTM onto a region that has some meaningful representation of the label characteristics, preparing the way for subsequent label prediction. Next is the CRF layer. CRF is a sequence probability model specifically designed to solve the optimal output label sequence under the condition of a given input sequence. It can explicitly model the dependencies between labels (for example, the end label of an entity cannot immediately follow the start label of another entity), thereby utilizing the global structural information of the label sequence to significantly improve the accuracy and consistency of the final sequence labeling.

[0055] In related technologies, web page content can be understood through manual reading, and then information can be manually extracted and a knowledge graph can be constructed to achieve information extraction of unstructured data on web pages, and effectively extract and utilize these unstructured data.

[0056] However, manual information extraction methods are not only time-consuming and difficult to adapt to the rapid growth of data volume, but also affect the processing results due to individual differences and subjective judgments, making it difficult to ensure the consistency of processing results and unable to achieve real-time updates and responses, resulting in low efficiency and accuracy in information extraction of unstructured data on web pages.

[0057] Based on this, the present application proposes an information extraction method to solve the above problems. The information extraction method provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0058] Figure 1 The flowchart of an information extraction method according to an exemplary embodiment is applied to a controller. The information extraction method includes the following steps.

[0059] In step S11 , the hypertext markup language HTML data of the target web page is obtained, and unstructured data in the HTML data is extracted.

[0060] HTML data is the basic code of a web page and contains all the elements that make up a web page, such as text, images, links, tables, etc. In this step, a web crawler or other data collection technology can be used to send a request to the target web page according to a network protocol (such as HTTP / HTTPS) to obtain its HTML data.

[0061] For example, in a Java development environment, you can use an HTTP client library such as Apache HttpClient or OkHttp to implement this process. First, you create an HTTP client instance and configure a request (such as HttpGet) that specifies the target web page URL. Then, by executing the request and obtaining a response, you can convert the response into a string for storage and processing, namely HTML data.

[0062] After obtaining HTML data, you can use HTML parsing tools or algorithms to parse it and convert it into a structured representation that is easy for programs to understand and manipulate. A common example is the DOM tree. By analyzing the structure and tags of the DOM tree, you can filter and extract unstructured data from the HTML data, including but not limited to text information such as news articles and product descriptions. This extracted unstructured data forms the basis for subsequent processing and analysis.

[0063] Among them, before converting HTML data into structured representation, preprocessing and data cleaning operations can be performed. Through these operations, useless information that interferes with analysis in the web page can be removed, such as advertising modules, navigation menus, irrelevant codes, and irrelevant links to other pages. Noise can be retained and focused on the core content of the web page, such as article titles, body paragraphs, publication dates, author information and other key information, thereby improving the accuracy and efficiency of subsequent processing.

[0064] In step S12, natural language processing is performed on the unstructured data to construct a knowledge graph, where the nodes in the knowledge graph represent entities in the unstructured data.

[0065] In this step, unstructured data is converted into structured knowledge representation and eventually built into a knowledge graph. In the knowledge graph, nodes are used to represent entities identified from the unstructured data.

[0066] First, a series of basic NLP processing can be performed on unstructured data, including but not limited to word segmentation, part-of-speech tagging, named entity recognition, etc. For example, the specific steps may include:

[0067] Word segmentation: splitting a continuous stream of text in unstructured data into independent vocabulary units;

[0068] Part-of-speech tagging: marking the part of speech (such as noun, verb, adjective, etc.) of each lexical unit after segmentation, which helps to understand the grammatical role of the word in the sentence;

[0069] Dependency Syntax Analysis: In-depth analysis of the grammatical dependencies between words, building a sentence dependency syntax tree, and revealing the deep connections between words;

[0070] Named entity recognition: Accurately identify entities with specific meanings from text, such as names of people (such as "Zhang San"), place names (such as "Beijing"), organization names (such as "XX Company"), etc. These identified entities will become the basic building blocks of the knowledge graph.

[0071] Then, based on these identified entities and the semantic relationships between them, a knowledge graph can be constructed to intuitively display the associations between data in a graphical manner. The knowledge graph includes nodes and edges. Each node represents an entity. For example, "XX Company" and "Zhang San" will both be nodes in the knowledge graph. By analyzing the sentences in the text that describe the interactions between entities, the semantic relationships between entities can be extracted. For example, the text may mention that "Zhang San once worked at XX Company", which indicates that there is a semantic relationship of "formerly worked at" between "Zhang San" and "XX Company". These semantic relationships are represented as edges connecting nodes in the knowledge graph.

[0072] In this way, entities and their semantic relationships in unstructured text are transformed into structured nodes and edges, ultimately forming a networked knowledge structure, the knowledge graph. The knowledge graph not only stores entity information but, more importantly, reveals the complex connections between entities, laying the foundation for subsequent knowledge reasoning.

[0073] In step S13, knowledge reasoning is performed on the knowledge graph based on the graph convolutional network to determine the implicit relationship between entities, and the knowledge graph is optimized based on the implicit relationship to obtain the information extraction result of the target web page.

[0074] While the knowledge graph structures entities and their semantic relationships, it may still contain implicit relationships that are not directly expressed. In this step, the knowledge graph is further mined and optimized, and graph convolutional networks are applied for knowledge reasoning to discover implicit relationships between entities. Based on these findings, the knowledge graph is optimized, ultimately generating high-quality information extraction results for the target web page.

[0075] Among them, the Graph Convolutional Network (GCN) is a deep learning model suitable for processing graph-structured data. By performing convolution operations on nodes and edges in a knowledge graph, it learns the feature information between nodes and infers potential, indirect, implicit relationships between entities, thereby deeply mining the potential semantic associations in the knowledge graph and expanding the information content of the knowledge graph. For example, even if the knowledge graph does not directly record that "XX Company" is a "consumer electronics manufacturer", if "XX Company" is associated with multiple "consumer electronics products" (such as "Product A" and "Product B"), the GCN may be able to infer this implicit classification relationship. Or, if it is known that "Zhang San" works at "a certain technology company" and "Li Si" also works at "a certain technology company", although the graph does not directly indicate the relationship between "Zhang San" and "Li Si", through the GCN's analysis and learning of node features, it may be possible to infer that they have a "colleague" relationship.

[0076] The working mechanism of a graph convolutional network is to take a node feature matrix (X) and an adjacency matrix (A, representing the edges between nodes) as input. It aggregates the features of neighboring nodes through the adjacency matrix to update the representation of the current node, and uses an activation function to introduce nonlinear transformations to enhance the model's expressive power. The activation function can use ReLU (Rectified Linear Unit). By repeating this process, the graph convolutional network can learn the high-order features of the nodes and ultimately output results that can be used for node classification or graph-level prediction. These results are the inferred implicit relationships or entity attributes.

[0077] When applying graph convolutional networks for reasoning, homomorphic encryption technology can be used to encrypt the input feature matrix X and adjacency matrix A, allowing the graph convolutional network to directly calculate the ciphertext data without decryption, thereby ensuring the data privacy and security of the reasoning process. In addition, SMPC (Secure Multi-Party Computation) or differential privacy technology can also be used to protect the graph convolutional network parameters and reasoning results from being leaked. In addition, the reasoning task can be distributed and deployed on multiple participants to prevent any single device from accessing the complete data set and graph convolutional network, thereby dispersing the risk of data leakage.

[0078] After identifying the implicit relationships between entities, these relationships can be added to the knowledge graph, optimizing and improving it. This supplements previously missed information, corrects potential errors, and enhances the completeness and accuracy of the knowledge graph. Ultimately, the optimized knowledge graph is the information extraction result of the target webpage, which not only contains information directly extracted from the original text but also incorporates the deeper knowledge mined by the deep learning model.

[0079] In one implementation, in step S11, extracting unstructured data from HTML data includes:

[0080] Parse HTML data and convert it into a Document Object Model (DOM) tree;

[0081] Traverse the DOM tree to perform structural analysis and determine the web page elements corresponding to the unstructured data in the HTML data;

[0082] Extract unstructured data from web page elements.

[0083] like Figure 2 The figure below is a schematic diagram of the HTML data parsing process in this embodiment. As can be seen from the above, HTML data is essentially a text code composed of a series of tags, which is difficult to process directly. Therefore, after obtaining the HTML data, you can first initialize the parsing tool and parse the obtained HTML data with the help of a specialized parsing tool (such as Jsoup, Beautiful Soup, etc.).

[0084] Specifically, HTML data can be converted into a tree structure, namely a DOM tree, based on the nesting level and sequence of HTML tags. In the DOM tree, each HTML tag corresponds to a node in the tree. Figure 3 The following is a schematic diagram of the hierarchical structure of HTML data, where the tag is the root node, and the tags such as , are child nodes under the root node. The child nodes of the tag include but are not limited to: specific text paragraphs ,picture ,sheet <h1>、< / h1> Tags such as <h2>; The sub-nodes of the tag include but are not limited to: <title> Tags, used to define external style content< / title> <link> Tags, used to define internal style content <style>标签、用于嵌入JavaScript代码或引用外部脚本文件的<script>标签。节点之间通过父子、兄弟等关系相互连接,形成一个完整的树状体系。通过这种转换,原本杂乱的HTML数据变得结构化,便于后续对其中特定元素进行定位和操作。

[0085] 将HTML数据转换为DOM树后,可以对DOM树进行遍历,遍历的方式通常有深度优先遍历或广度优先遍历等。在遍历过程中,可以分析每个节点的标签名称、属性以及节点之间的语义关系。

[0086] 通过研究不同网页的结构规律,并结合具体的业务需求,可以确定哪些节点对应的网页元素包含非结构化数据。例如,新闻正文通常会包含在标签或者特定类名、ID的标签中;标题可能在<h1>标签或者具有特定标识的元素里。对于一些结构更为复杂的目标网页,可能还需要分析多层嵌套的节点语义关系,比如在包含表格的目标网页中,表格所在的标签及其子节点(行)、(单元格)可能需要进一步判断和筛选,以确定其是否为所需的非结构化数据来源。通过这样的结构分析,可以精准定位到包含非结构化数据的网页元素。

[0087] 定位到包含非结构化数据的网页元素后,便可以从这些网页元素中提取出具体的非结构化数据。对于文本类型的非结构化数据,直接获取节点中的文本内容即可。例如,对于标签内的新闻正文,解析工具可以直接提取标签之间的文本字符串。如果是其他形式的非结构化数据,如图片链接、视频链接等,则需要提取节点的特定属性值。例如,对于标签,通过获取其src属性值,得到图片的链接地址;对于标签,获取href属性值,得到链接指向的目标地址。

[0088] 提取出非结构化数据后,还可以进行一些初步的数据清洗和格式化处理,以得到数据输出,比如去除多余的空白字符、特殊符号,或者对编码格式进行转换,以确保提取出的非结构化数据符合后续处理的要求,为后续构建知识图谱等操作提供准确、可用的数据基础。

[0089] 一种实现方式中,从网页元素中提取出非结构化数据,包括:

[0090] 在网页元素中包括表格的情况下,根据预设拆分策略,遍历网页元素中的表格,提取出表格数据;

[0091] 创建表格数据对应的数据存储单元,并将表格数据存储至数据存储单元中,得到非结构化数据。

[0092] 在本实施例中,当通过对DOM树的分析,确定网页元素中存在表格(标签)时,处理流程会进入专门针对表格的数据提取环节。首先,需要明确预设拆分策略,预设拆分策略是根据具体业务需求和处理目标提前设定的规则,常见的包括按行拆分、按列拆分,或根据特定单元格内容拆分。

[0093] 其中,按行拆分是指以表格的"行”(标签)为基本单位,从表格的第一行开始,逐行提取每个单元格(或标签)内的表格数据。例如,在一个人员信息表中,每一行对应一个人员的各项信息,按行拆分可以将每个人员的信息完整提取出来,作为一个独立的数据单元。

[0094] 按列拆分是指以表格的"列”为基本单位,从表格的第一列开始,逐列提取每个单元格内的表格数据。比如,在统计销售数据的表格中,按列拆分能将产品名称、销售数量、销售额等不同属性的数据分别归集,便于后续对特定属性数据进行分析。

[0095] 根据特定单元格内容拆分是指先定位到包含关键信息(如姓名、编号、类别等)的单元格,然后依据这些单元格内容的不同取值进行拆分。例如,在一个包含多个部门员工信息的表格中,根据"部门名称”单元格的内容,将属于同一部门的员工信息拆分到一起,形成多个子表格或数据块。

[0096] 在确定预设拆分策略后,可以利用解析工具对表格对应的DOM树节点进行遍历。在遍历过程中,依据预设拆分策略提取相应的表格数据。对于跨行或跨列的复杂单元格,需要特别处理,确保数据提取的完整性和准确性。例如,对于合并单元格,需要将合并区域的数据按照逻辑分配到拆分后的数据单元中。

[0097] 提取出表格数据后,可以为其创建合适的数据存储单元,以便进行后续处理和管理。数据存储单元的形式取决于具体需求和后续应用场景。比如,如果需要保留表格的结构特征,可创建二维数组、列表嵌套等数据结构。例如,使用二维数组[[单元格1,单元格2,…],[单元格1,单元格2,…],…]来存储按行拆分后的表格数据,数组的每个子数组对应表格的一行,这种方式能够清晰地反映表格的行列结构,适用于需要对表格数据进行行列操作的场景。

[0098] 如果不需要保留表格格式和结构,可将表格数据存储到更灵活的数据结构中,如对象列表或字典列表。例如,将每一行数据转换为一个对象或字典,其属性名对应列标题,属性值对应单元格内容,如{"姓名":"张三","年龄":25,"部门":"技术部"}。这种方式便于对表格数据进行灵活的查询、筛选和分析。

[0099] 进而,将提取到的表格数据按照选定的数据存储单元格式进行整理和存储,最终得到的存储结果即为从网页表格元素中提取出的非结构化数据,用于构建知识图谱。

[0100] 一种实现方式中,在步骤S12中,对非结构化数据进行自然语言处理,构建知识图谱,包括:

[0101] 对非结构化数据进行自然语言处理,提取非结构化数据中的实体、实体间的语义关系以及实体的属性信息;

[0102] 根据实体、实体间的关系以及实体的属性信息,构建知识图谱,其中,知识图谱中的节点之间的边表示对应的实体间的语义关系,属性信息用于表示节点或边的属性。

[0103] 可以理解,非结构化数据通常以自由文本形式存在,缺乏明确的数据结构,因此需借助自然语言处理技术对其进行剖析,提取非结构化数据中的实体、实体间的语义关系以及实体的属性信息。

[0104] 其中,文本中具有特定意义的实体包括人名(如"张三”)、地名(如"北京”)、组织机构名(如"XX公司”)、产品名(如"产品A”)等。实体之间的语义关系包是指实体之间的关联,例如,通过分析句子"张三创办了XX公司”,可确定"张三”与"XX公司”之间存在"创办”的关系。每个实体通常具有一系列属性来描述其特征,如"XX公司”的属性可能包括"成立时间”"总部地点”等。

[0105] 进而,在获取实体、关系和属性信息后,可以将这些元素整合构建知识图谱。具体地,将提取出的实体作为知识图谱中的节点,每个节点代表一个现实世界中的对象或概念。例如,"张三”"XX公司”分别作为独立的节点存在于知识图谱中。然后,以实体间的语义关系为依据,在对应的实体节点之间建立边,边的方向和标签表示关系的类型和方向,如从"张三”节点指向"XX公司”节点的边,标签为"创办”,清晰展示了两者之间的关系。进而,可以将提取的属性信息作为节点或边的属性进行添加。例如,"XX公司”节点可以添加"成立时间:1999年”"总部地点:北京”等属性;而边也可以具有属性,如"创办”关系边可添加"创办方式:独资”等属性。通过这种方式,知识图谱不仅呈现了实体间的关系,还丰富了每个节点和边的细节信息。

[0106] 通过以上步骤,原本无序的非结构化数据被转化为结构化的知识图谱,以图形化的方式直观地展示了数据中的知识关联。

[0107] 一种实现方式中,对非结构化数据进行自然语言处理,提取非结构化数据中的实体、实体间的语义关系以及实体的属性信息,包括:

[0108] 对非结构化数据进行命名实体识别,提取非结构化数据中的实体,并识别出与实体相关的属性信息;

[0109] 对非结构化数据进行语义识别,得到实体之间的语义关系。

[0110] 在本实施例中,命名实体识别是自然语言处理中的一项基础且关键的任务,针对非结构化数据,可以利用命名实体识别算法或模型对文本进行逐词、逐句分析,旨在从非结构化数据中找出具有特定意义的实体,并分类为预定义的类别(如人名、地名、组织机构名、时间、产品名等)。命名实体识别的实现可以基于规则、统计模型或深度学习算法(例如Bi-LSTM-CRF模型),具体不做限定。例如,如图4所示,对于新闻文本"2024年7月30日,某科技公司在北京举办了产品发布会,CEO张三发布了新款香蕉手机”,命名实体识别模型可以快速定位并提取出"某科技公司”(组织机构名)、"北京”(地名)、"香蕉手机”(产品名)等实体。

[0111] 属性信息是对实体特征、性质或状态的描述。在识别实体的过程中,可以同时提取与实体相关的属性信息。可以采用基于规则的方法(如人工定义规则和模式匹配)、基于统计学习的方法(如从文本中提取词性、词频、上下文特征)或基于深度学习的方法(如利用大规模预训练模型进行微调),分析实体前后的词汇、语法结构,判断哪些词汇是描述该实体的属性信息。例如,上述例子中,"新款”是"香蕉手机”的属性信息,描述了产品的特征;"北京”可以看作是"产品发布会”发生地的属性。

[0112] 在识别出实体后,可以进一步分析文本语义,语义识别旨在理解文本中词语、句子的含义,从而挖掘出实体之间存在的语义关联。例如,通过分析句子"张三创办了XX公司”,利用语义关系提取技术可确定"张三”与"XX公司”之间存在"创办”的关系。对于复杂的文本,可能存在多个实体及多种关系交织的情况,需综合运用多种方法,准确梳理出实体间的语义关系。

[0113] 其中,在对非结构化数据进行语义分析之前,可以对非结构化数据进行预处理,包括分词、词性标注、句法分析等步骤。例如,将句子"XX公司收购了YY公司”进行分词为"XX公司 / 收购 / 了 / YY公司”,并标注词性("XX公司”:名词,"收购”:动词等),通过句法分析确定句子的结构(主谓宾结构)。这些预处理为后续语义关系的提取提供基础。

[0114] 对预处理后的非结构化数据,可以运用模式匹配、依存句法分析或深度学习模型来提取实体间的语义关系。其中,模式匹配是通过预先定义好的关系模板(如"[组织机构名]收购[组织机构名 / 产品名]”),在文本中寻找符合模板的语句,从而确定实体间的关系;依存句法分析通过分析句子中词语之间的依存关系,找出与实体相关的动词或关系词(例如通过分析"收购”这个动词),进而确定实体间的关系;深度学习模型则通过在大量标注了关系的文本数据上进行训练,自动学习文本中实体与关系之间的语义模式,从而准确识别出各种语义关系。

[0115] 举例来说,可以采用基于注意力机制的深度学习模型识别实体之间的语义关系,基于注意力机制的深度学习模型的核心思想是通过动态分配权重,聚焦于输入序列中对当前任务最重要的部分,从而提升模型在长序列依赖建模、信息聚合和特征提取等方面的能力,实体之间的语义关系可以以三元组的形式表示,即(主体实体,关系类型,客体实体)。

[0116] 其中,在基于注意力机制的深度学习模型中,模型为输入序列中的每个元素计算一个权重,权重越高表示该元素对当前任务越重要。而且,注意力机制允许模型在处理每个输出时,动态地关注输入序列的不同部分,而非固定地使用整个输入。进而,可以通过加权求和的方式,将输入序列中重要的信息聚合到输出中。

[0117] 基于注意力机制的深度学习模型通常包括以下组件:

[0118] 编码器(Encoder):将输入序列(如文本、图像)编码为隐藏表示(hiddenrepresentations)。

[0119] 注意力层(Attention Layer):计算输入序列中每个元素的注意力权重,并生成加权后的上下文向量(context vector)。

[0120] 解码器(Decoder):使用上下文向量生成输出序列(如翻译文本、分类标签)。

[0121] 注意力机制的计算过程包括:

[0122] 查询(Query)、键(Key)、值(Value):输入序列中的每个元素被映射为查询(Q)、键(K)和值(V)。例如,在自然语言处理中,Q、K、V通常由输入的词向量通过线性变换得到。

[0123] 相似度计算:计算查询(Q)与键(K)之间的相似度(如点积、余弦相似度)。

[0124] 注意力权重:对相似度进行归一化,得到注意力权重。

[0125] 加权求和:使用注意力权重对值(V)进行加权求和,得到上下文向量。

[0126] 通过上述两个步骤,能够系统地从非结构化数据中提取出实体、属性信息以及实体间的语义关系,为后续构建知识图谱提供了数据支持。

[0127] 一种实现方式中,在步骤S13中,基于图卷积网络对知识图谱进行知识推理,确定实体间的隐含关系,并基于隐含关系对知识图谱进行优化,得到目标网页的信息抽取结果,包括:

[0128] 对知识图谱进行去重处理、消歧处理和逻辑校验,得到优化后的知识图谱;

[0129] 基于图卷积网络对优化后的知识图谱进行知识推理,挖掘实体间的隐含关系;

[0130] 将隐含关系补充至优化后的知识图谱中,得到目标网页的信息抽取结果。

[0131] 可以理解,在知识图谱构建过程中,由于数据来源多样或提取过程存在误差,可能会引入冗余、歧义或逻辑矛盾的信息,因此,在应用图卷积网络进行知识推理之前,可以先进行基础优化,其中包括:

[0132] 去重处理:知识图谱中可能存在重复的实体、关系或属性信息。例如,不同数据源可能产生关于同一实体(如"XX公司”)的多个相似节点,或者同一种关系(如"投资”)被多次重复记录。去重处理通过对比节点的名称、属性,以及关系的起始节点、终止节点和类型等信息,识别并合并重复项,消除冗余数据,确保每个实体、语义关系在知识图谱中仅以唯一的形式存在,使知识图谱结构更加简洁高效。

[0133] 消歧处理:同一名称的实体可能对应不同的现实对象,从而产生歧义。例如,"XX”既可以指代某种水果,也可以指代XX公司。消歧处理通过分析实体所在的上下文信息(如与之关联的其他实体、描述性属性等),结合外部知识库(如百科知识),明确每个实体的准确含义。例如,若"XX”节点关联了"手机”"操作系统”等属性,则可判定其为XX公司;若关联"水果种类”"口感”等属性,则可判定为水果。通过这种方式消除歧义,确保知识图谱中每个实体指代清晰,避免推理过程中的混淆。

[0134] 逻辑校验:检查知识图谱中实体之间的关系、属性之间的关联是否符合逻辑规则。例如,在人物关系中,若存在"A是B的父亲”且"B是A的父亲”这样相互矛盾的关系;或者在时间属性上,一个事件的发生时间晚于其结束时间,这些都是逻辑错误。逻辑校验通过预先设定的逻辑规则(如父子关系的单向性、时间先后顺序等),对知识图谱进行全面检查,修正或删除不符合逻辑的关系和属性,保证知识图谱的逻辑一致性和合理性,为后续推理提供可靠基础。

[0135] 由前述可知,图卷积网络能够推断出优化后的知识图谱中实体之间尚未明确表示的隐含关系。在挖掘出实体间的隐含关系后,可以将这些隐含关系融入知识图谱,进一步完善知识图谱。具体的,可以将推理得到的隐含关系以边的形式添加到优化后的知识图谱中,连接对应的实体节点,并标注关系类型。例如,将"张三”和"李四”通过"同事”关系边连接起来,同时可以为这条边添加相关属性(如"同事关系起始时间”),丰富关系的细节信息。通过这种方式,知识图谱的结构更加完整,涵盖了更多的语义关联,从而得到了经过优化和推理扩展的目标网页信息抽取结果。

[0136] 经过隐含关系补充后的知识图谱,包含了从目标网页数据中提取并经过推理扩展的全面信息,形成了最终的信息抽取结果。该结果可以为知识查询、智能问答、数据分析等多种应用提供结构化的知识支持,例如在智能问答系统中,能够基于完整的知识图谱更准确地回答用户关于实体关系的问题,实现对网页信息的深度利用和价值挖掘。

[0137] 一种实现方式中,在步骤S13之后,还包括:

[0138] 将信息抽取结果存储到数据库中,并在数据库中创建信息抽取结果的索引,索引用于对信息抽取结果进行查询。

[0139] 在本实施例中,将信息抽取结果存储到数据库中,以实现信息抽取结果的持久化保存,避免数据丢失,同时方便后续的管理和应用。

[0140] 首先,可以根据知识图谱的特点和应用需求,选择与之适配的数据库类型,确保知识图谱的结构和语义信息能够被准确、高效地存储和检索。其中,数据库类型可以包括但不限于图数据库(如Neo4j),它能够原生支持图结构数据的存储和查询,非常适合存储知识图谱,可高效处理节点和边的复杂关系;也可以选择关系型数据库(如MySQL)或非关系型数据库(如MongoDB),此时需要将知识图谱的数据进行适当的转换,例如将节点和关系分别存储为表中的记录,通过外键等方式建立关联。

[0141] 确定数据库类型后,可以设计数据映射规则,将知识图谱中的实体、关系和属性准确地存储到数据库中。对于图数据库,节点直接对应数据库中的节点记录,边对应关系记录,属性作为节点或边的属性字段存储;若使用关系型数据库,可能需要创建多个表,如实体表存储节点信息、关系表存储边的信息,通过唯一标识建立表间关联。在存储过程中,需要确保数据的完整性和准确性,例如避免实体属性的丢失或关系连接错误。

[0142] 进而,可以使用数据库提供的索引创建功能,按照选定的索引字段生成索引。其中,创建索引是为了加快对信息抽取结果的查询速度,提升数据检索的效率,尤其当数据库中的数据量较大时,索引的作用更为显著。具体地,可以根据实际的查询需求,选择频繁用于查询条件的属性或关系作为索引字段。例如,如果经常根据实体名称(如人名、公司名)查询相关信息,那么可以将实体名称字段设置为索引;若经常查询特定关系类型(如"合作”"投资”)的所有关系实例,则可对关系类型字段创建索引。此外,复合索引也是常用方式,例如同时基于实体名称和关系类型创建索引,以满足更复杂的查询条件。

[0143] 通过将信息抽取结果存储到数据库并创建索引,不仅实现了数据的安全存储,还为后续诸如知识检索、数据分析、智能应用等场景提供了高效的数据访问支持。

[0144] 图5是根据一示例性实施例示出的一种信息抽取方法的实例示意图,其中包括如下步骤:

[0145] 数据获取:开发高效的网络爬虫,自动抓取目标网页的HTML数据。

[0146] 数据预处理与提取:清洗HTML数据,利用DOM树或XPath等技术,构建出网页元素的层次结构和关联关系,定义数据抽取规则,根据抽取规则对目标网页中的非结构化数据进行提取和处理。

[0147] 数据映射与转换:定义相应的数据结构模板(字段名称、数据类型、约束条件等),用于指导非结构化数据的转换和输出。通过去除重复数据、处理缺失值、转换数据类型等操作,确保输出的结构化数据具有高质量和一致性。

[0148] 文本解析:采用自然语言处理技术对非结构化数据进行分词、词性标注和依存句法分析,将连续的文本分割成独立的词汇单元,词性标注模块为每个词汇标注词性信息,依存句法分析模块则分析词汇之间的语法关系。

[0149] 信息提取:包含实体识别、关系抽取、属性抽取,并通过实体链接将识别的实体与现有知识图谱中的实体进行匹配,解决实体消岐问题。

[0150] 知识图谱构建与优化:整合抽取的实体、关系和属性,构建知识图谱,对构建的知识图谱进行去重、合并和属性标准化等操作,提高知识图谱的质量和可用性。

[0151] 知识图谱存储:将知识图谱存储至图数据库或其他适合的存储系统中。

[0152] 其中,若任一步骤执行失败,则记录异常并结束信息抽取流程。

[0153] 由以上可见,本申请的实施例提供的技术方案,本申请通过自动化流程获取并处理目标网页的HTML数据,能够大幅提升信息抽取的效率,有效应对海量数据的挑战,而且,基于自然语言处理和图卷积网络进行知识图谱构建与推理,不仅实现了处理过程的标准化,确保了结果的一致性,减少了主观判断的影响,还能深入挖掘并确定实体间的隐含关系,从而显著提升了信息抽取的准确性和深度,此外,这种自动化、智能化的方法为实时更新和响应网页数据变化提供了可能,克服了人工方式的实时性限制。

[0154] 本申请实施例提供的信息抽取方法,执行主体可以为信息抽取装置。本申请实施例中以信息抽取装置执行信息抽取的方法为例,说明本申请实施例提供的信息抽取方法的装置。

[0155] 图6是根据一示例性实施例示出的一种信息抽取装置框图,包括:

[0156] 获取模块301,用于获取目标网页的超文本标记语言HTML数据,并提取出所述HTML数据中的非结构化数据;

[0157] 构建模块302,用于对所述非结构化数据进行自然语言处理,构建知识图谱,其中,所述知识图谱中的节点表示所述非结构化数据中的实体;

[0158] 优化模块303,用于基于图卷积网络对所述知识图谱进行知识推理,确定所述实体间的隐含关系,并基于所述隐含关系对所述知识图谱进行优化,得到所述目标网页的信息抽取结果。

[0159] 一种实现方式中,所述获取模块301,具体用于:

[0160] 解析所述HTML数据,将所述HTML数据转换为文档对象模型DOM树;

[0161] 遍历所述DOM树进行结构分析,确定所述HTML数据中的非结构化数据对应的网页元素;

[0162] 从所述网页元素中提取出所述非结构化数据。

[0163] 一种实现方式中,所述获取模块301,具体用于:

[0164] 在所述网页元素中包括表格的情况下,根据预设拆分策略,遍历所述网页元素中的表格,提取出表格数据;

[0165] 创建所述表格数据对应的数据存储单元,并将所述表格数据存储至所述数据存储单元中,得到所述非结构化数据。

[0166] 一种实现方式中,所述构建模块302,具体用于:

[0167] 对所述非结构化数据进行自然语言处理,提取所述非结构化数据中的实体、所述实体间的语义关系以及所述实体的属性信息;

[0168] 根据所述实体、所述实体间的关系以及所述实体的属性信息,构建知识图谱,其中,所述知识图谱中的节点之间的边表示对应的所述实体间的语义关系,所述属性信息用于表示所述节点或所述边的属性。

[0169] 一种实现方式中,所述构建模块302,具体用于:

[0170] 对所述非结构化数据进行命名实体识别,提取所述非结构化数据中的实体,并识别出与所述实体相关的属性信息;

[0171] 对所述非结构化数据进行语义识别,得到所述实体之间的语义关系。

[0172] 一种实现方式中,所述优化模块303,具体用于:

[0173] 对所述知识图谱进行去重处理、消歧处理和逻辑校验,得到优化后的知识图谱;

[0174] 基于图卷积网络对所述优化后的知识图谱进行知识推理,挖掘所述实体间的隐含关系;

[0175] 将所述隐含关系补充至所述优化后的知识图谱中,得到所述目标网页的信息抽取结果。

[0176] 一种实现方式中,所述装置还包括:

[0177] 存储模块,用于将所述信息抽取结果存储到数据库中,并在所述数据库中创建所述信息抽取结果的索引,所述索引用于对所述信息抽取结果进行查询。

[0178] 由以上可见,本申请的实施例提供的技术方案,本申请通过自动化流程获取并处理目标网页的HTML数据,能够大幅提升信息抽取的效率,有效应对海量数据的挑战,而且,基于自然语言处理和图卷积网络进行知识图谱构建与推理,不仅实现了处理过程的标准化,确保了结果的一致性,减少了主观判断的影响,还能深入挖掘并确定实体间的隐含关系,从而显著提升了信息抽取的准确性和深度,此外,这种自动化、智能化的方法为实时更新和响应网页数据变化提供了可能,克服了人工方式的实时性限制。

[0179] 本申请实施例提供的信息抽取方法,执行主体可以为终端接入终端。本申请实施例中以终端接入终端执行终端接入的方法为例,说明本申请实施例提供的信息抽取方法的装置。

[0180] 本申请实施例中的信息抽取装置可以是电子设备,也可以是电子设备中的部件,例如集成电路或芯片。该电子设备可以是终端,也可以为除终端之外的其他设备。示例性的,电子设备可以为手机、平板电脑、笔记本电脑、掌上电脑、车载电子设备、移动上网装置(Mobile Internet Device,MID)、增强现实(augmented reality,AR) / 虚拟现实(virtualreality,VR)设备、机器人、可穿戴设备、超级移动个人计算机(ultra-mobile personalcomputer,UMPC)、上网本或者个人数字助理(personal digital assistant,PDA)等,还可以为服务器、网络附属存储器(Network Attached Storage,NAS)、个人计算机(personalcomputer,PC)、电视机(television,TV)、柜员机或者自助机等,本申请实施例不作具体限定。

[0181] 本申请实施例提供的信息抽取装置能够实现图1至图5的方法实施例实现的各个过程,为避免重复,这里不再赘述。

[0182] 可选地,如图7所示,本申请实施例还提供一种电子设备500,包括处理器501和存储器502,存储器502上存储有可在所述处理器501上运行的程序或指令,该程序或指令被处理器501执行时实现上述信息抽取方法实施例的各个步骤,且能达到相同的技术效果,为避免重复,这里不再赘述。

[0183] 需要说明的是,本申请实施例中的电子设备包括上述所述的移动电子设备和非移动电子设备。

[0184] 图8为实现本申请实施例的一种电子设备的硬件结构示意图。

[0185] 该电子设备1000包括但不限于:射频单元1001、网络模块1002、音频输出单元1003、输入单元1004、传感器1005、显示单元1006、用户输入单元1007、接口单元1008、存储器1009、以及处理器1010等部件。

[0186] 本领域技术人员可以理解,电子设备1000还可以包括给各个部件供电的电源(比如电池),电源可以通过电源管理系统与处理器1010逻辑相连,从而通过电源管理系统实现管理充电、放电、以及功耗管理等功能。图8中示出的电子设备结构并不构成对电子设备的限定,电子设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置,在此不再赘述。

[0187] 由以上可见,本申请的实施例提供的技术方案,本申请通过自动化流程获取并处理目标网页的HTML数据,能够大幅提升信息抽取的效率,有效应对海量数据的挑战,而且,基于自然语言处理和图卷积网络进行知识图谱构建与推理,不仅实现了处理过程的标准化,确保了结果的一致性,减少了主观判断的影响,还能深入挖掘并确定实体间的隐含关系,从而显著提升了信息抽取的准确性和深度,此外,这种自动化、智能化的方法为实时更新和响应网页数据变化提供了可能,克服了人工方式的实时性限制。

[0188] 应理解的是,本申请实施例中,输入单元1004可以包括图形处理器(GraphicsProcessing Unit,GPU)10041和麦克风10042,图形处理器10041对在视频捕获模式或图像捕获模式中由图像捕获装置(如摄像头)获得的静态图片或视频的图像数据进行处理。显示单元1006可包括显示面板10061,可以采用液晶显示器、有机发光二极管等形式来配置显示面板10061。用户输入单元1007包括触控面板10071以及其他输入设备10072中的至少一种。触控面板10071,也称为触摸屏。触控面板10071可包括触摸检测装置和触摸控制器两个部分。其他输入设备10072可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆,在此不再赘述。

[0189] 存储器1009可用于存储软件程序以及各种数据。存储器1009可主要包括存储程序或指令的第一存储区和存储数据的第二存储区,其中,第一存储区可存储操作系统、至少一个功能所需的应用程序或指令(比如声音播放功能、图像播放功能等)等。此外,存储器1009可以包括易失性存储器或非易失性存储器,或者,存储器1009可以包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synch link DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请实施例中的存储器109包括但不限于这些和任意其它适合类型的存储器。

[0190] 处理器1010可包括一个或多个处理单元;可选的,处理器1010集成应用处理器和调制解调处理器,其中,应用处理器主要处理涉及操作系统、用户界面和应用程序等的操作,调制解调处理器主要处理无线通信信号,如基带处理器。可以理解的是,上述调制解调处理器也可以不集成到处理器1010中。

[0191] 本申请实施例还提供一种可读存储介质,所述可读存储介质上存储有程序或指令,该程序或指令被处理器执行时实现上述信息抽取方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。

[0192] 其中,所述处理器为上述实施例中所述的电子设备中的处理器。所述可读存储介质,包括计算机可读存储介质,如计算机只读存储器ROM、随机存取存储器RAM、磁碟或者光盘等。

[0193] 本申请实施例另提供了一种芯片,所述芯片包括处理器和通信接口,所述通信接口和所述处理器耦合,所述处理器用于运行程序或指令,实现上述信息抽取方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。

[0194] 应理解,本申请实施例提到的芯片还可以称为系统级芯片、系统芯片、芯片系统或片上系统芯片等。

[0195] 本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如上述信息抽取方法实施例的各个过程,且能达到相同的技术效果,为避免重复,这里不再赘述。

[0196] 需要说明的是,在本文中,术语"包括”、"包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句"包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。此外,需要指出的是,本申请实施方式中的方法和装置的范围不限按示出或讨论的顺序来执行功能,还可包括根据所涉及的功能按基本同时的方式或按相反的顺序来执行功能,例如,可以按不同于所描述的次序来执行所描述的方法,并且还可以添加、省去、或组合各种步骤。另外,参照某些示例所描述的特征可在其他示例中被组合。

[0197] 通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到上述实施例方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分可以以计算机软件产品的形式体现出来,该计算机软件产品存储在一个存储介质(如ROM / RAM、磁碟、光盘)中,包括若干指令用以使得一台终端(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例所述的方法。

[0198] 上面结合附图对本申请的实施例进行了描述,但是本申请并不局限于上述的具体实施方式,上述的具体实施方式仅仅是示意性的,而不是限制性的,本领域的普通技术人员在本申请的启示下,在不脱离本申请宗旨和权利要求所保护的范围情况下,还可做出很多形式,均属于本申请的保护之内。< / style> < / h2>

Claims

1. An information extraction method, characterized in that: The method comprises: Obtaining Hypertext Markup Language (HTML) data of a target web page and extracting unstructured data from the HTML data; Performing natural language processing on the unstructured data to construct a knowledge graph, wherein nodes in the knowledge graph represent entities in the unstructured data; The knowledge graph is subjected to knowledge reasoning based on a graph convolutional network to determine the implicit relationships between the entities, and the knowledge graph is optimized based on the implicit relationships to obtain the information extraction result of the target web page.

2. The information extraction method according to claim 1, wherein: The extracting of unstructured data from the HTML data includes: Parsing the HTML data and converting the HTML data into a Document Object Model (DOM) tree; Traversing the DOM tree to perform structural analysis and determine the web page elements corresponding to the unstructured data in the HTML data; The unstructured data is extracted from the web page elements.

3. The information extraction method according to claim 2, characterized in that The extracting the unstructured data from the web page elements includes: In the case where the web page element includes a table, traversing the table in the web page element according to a preset splitting strategy to extract the table data; A data storage unit corresponding to the table data is created, and the table data is stored in the data storage unit to obtain the unstructured data.

4. The information extraction method according to claim 1, wherein: The performing natural language processing on the unstructured data to construct a knowledge graph includes: Performing natural language processing on the unstructured data to extract entities in the unstructured data, semantic relationships between the entities, and attribute information of the entities; A knowledge graph is constructed based on the entities, the relationships between the entities, and the attribute information of the entities, wherein the edges between the nodes in the knowledge graph represent the corresponding semantic relationships between the entities, and the attribute information is used to represent the attributes of the nodes or the edges.

5. The information extraction method according to claim 4, characterized in that The performing natural language processing on the unstructured data to extract entities in the unstructured data, semantic relationships between the entities, and attribute information of the entities includes: Performing named entity recognition on the unstructured data, extracting entities from the unstructured data, and identifying attribute information related to the entities; Semantic recognition is performed on the unstructured data to obtain semantic relationships between the entities.

6. The information extraction method according to claim 1, characterized in that The graph convolutional network-based knowledge reasoning on the knowledge graph, determining the implicit relationship between the entities, and optimizing the knowledge graph based on the implicit relationship to obtain the information extraction result of the target web page includes: Performing deduplication processing, disambiguation processing, and logical verification on the knowledge graph to obtain an optimized knowledge graph; Performing knowledge reasoning on the optimized knowledge graph based on a graph convolutional network to mine implicit relationships between the entities; The implicit relationship is added to the optimized knowledge graph to obtain the information extraction result of the target web page.

7. The information extraction method according to claim 1, characterized in that After optimizing the knowledge graph based on the implicit relationship to obtain the information extraction result of the target webpage, the method further includes: The information extraction result is stored in a database, and an index of the information extraction result is created in the database, where the index is used to query the information extraction result.

8. An information extraction device, characterized in that: The device comprises: An acquisition module, configured to acquire the Hypertext Markup Language (HTML) data of a target web page and extract unstructured data from the HTML data; A construction module, configured to perform natural language processing on the unstructured data to construct a knowledge graph, wherein nodes in the knowledge graph represent entities in the unstructured data; An optimization module is used to perform knowledge reasoning on the knowledge graph based on a graph convolutional network, determine the implicit relationship between the entities, and optimize the knowledge graph based on the implicit relationship to obtain the information extraction result of the target web page.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the information extraction method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the information extraction method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Knowledge graph construction method and device, electronic equipment and storage medium

    CN121279419A