Internet Online Common Sense Extraction Method, Device, Electronic Device and Storage Medium

By segmenting and co-referential dissolution of long texts, calculating the total score of global co-referentials and determining the co-referential set, the problem of insufficient accuracy of knowledge extraction in long texts is solved, and more efficient entity relationship extraction and knowledge extraction are achieved.

CN119783675BActive Publication Date: 2025-07-01BEI JING BDA NETWORK &INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272530.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-01
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The prior art lacks good long-distance modeling capabilities when processing long texts, resulting in insufficient accuracy and completeness of knowledge extraction.

Method used

By dividing the Internet documents into sub-documents, performing co-referential digestion and entity recognition, building a co-referential sparse matrix, calculating the total score of the total score of the total score of the total score of the total score of the total score of the total score of the total score of the total score of the total score of the total score of the total score of the global co-referential, determining the set of the global co-referential, and performing entity relationship extraction of the updated documents.

Benefits of technology

Overcome the problem of long-distance dependence in long texts, improve the ability to extract entity relationships, and improve the accuracy of knowledge extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783675B_ABST
    Figure CN119783675B_ABST
Patent Text Reader

Abstract

The present invention provides an Internet online common sense extraction method, device, electronic device and storage medium. By performing coreference resolution and entity recognition on sub-documents of Internet documents, a plurality of coreference sets and the entities existing in each coreference set are obtained, and a coreference sparse matrix corresponding to each coreference set of each sub-document is constructed. Furthermore, based on this, the coreference total score of two entity mentions in the sub-document is calculated. Based on the coreference total scores of two entity mentions in each sub-document, the global coreference total score of any two entity mentions in the Internet document is calculated. Then, the global coreference set is determined. Thus, based on the entity names of the entities in each global coreference set, the entity mentions in the corresponding global coreference set are replaced to obtain an updated document, and entity relationship extraction is performed on the updated document to obtain triples of the Internet document, and the knowledge base is updated based on the triples of the Internet document, which improves the entity relationship extraction ability of long texts and improves the knowledge extraction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to an Internet online common sense extraction method, device, electronic device, and storage medium. Background Art

[0002] Automatically extracting commonsense knowledge from the Internet and constructing a knowledge base is one of the important tasks in the field of artificial intelligence. The core of this process includes identifying key entities and the relationships between them from unstructured text, thereby generating triples consisting of a head entity, an entity relationship, and a tail entity. Among them, in the initial stage of knowledge extraction, entity recognition is used to locate and classify key terms in the text, such as personal names, place names, or organization names. This task requires the model to be able to understand the context and distinguish which phrases are real entities. Then, relation extraction technology is used to identify the associations between these entities, such as "so-and-so founded a certain company" or "a certain event occurred in a certain place". These structured data provide an important basis for question answering systems, information retrieval, and decision support.

[0003] However, there are a large number of long texts in Internet documents, and long texts bring many challenges to current natural language processing tasks. In short texts, the recognition of entities and relationships is relatively easy because the information is more concentrated and the connections between sentences are relatively simple. However, for long texts, the problem becomes much more complex. The content of long texts often consists of multiple topics, and its description methods often use pronouns or synonymous phrases to repeatedly mention the same entity, resulting in some key information being scattered across multiple paragraphs. Currently, due to the lack of good long-distance dependence modeling ability, the model cannot associate semantic information that is far apart, affecting the accuracy and completeness of knowledge extraction. Summary of the Invention

[0004] The present invention provides an Internet online common sense extraction method, device, electronic device, and storage medium to solve the defect of poor accuracy of knowledge extraction for long texts in the prior art.

[0005] The present invention provides an Internet online common sense extraction method, including:

[0006] Segmenting an Internet document to obtain a plurality of sub-documents;

[0007] Performing coreference resolution and entity recognition on any one of the sub-documents to obtain a plurality of coreference sets and the entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of the any one of the sub-documents; wherein, the elements in the coreference sparse matrix corresponding to any one coreference set represent the coreference scores of two entity mentions in the any one of the sub-documents with respect to the any one coreference set.

[0008] Based on the co-reference sparse matrix corresponding to each co-reference set of any sub-document, calculate the co-reference total score of two entity mentions in the any sub-document. Based on the co-reference total scores of two entity mentions in each sub-document, calculate the global co-reference total score of any two entity mentions in the Internet document, and determine the global co-reference set based on the global co-reference total score of any two entity mentions in the Internet document;

[0009] Based on the entity names of the entities in each global co-reference set, replace the entity mentions in the corresponding global co-reference set to obtain an updated document, perform entity relationship extraction on the updated document to obtain the triples of the Internet document, and update the knowledge base based on the triples of the Internet document.

[0010] According to an Internet online common sense extraction method provided by the present invention, the calculating the global co-reference total score of any two entity mentions in the Internet document based on the co-reference total scores of two entity mentions in each sub-document includes:

[0011] Construct a non-co-reference sparse matrix corresponding to each co-reference set of each sub-document, and calculate the non-co-reference total score of two entity mentions in each sub-document based on the non-co-reference sparse matrix corresponding to each co-reference set of each sub-document; wherein, the elements in the non-co-reference sparse matrix corresponding to any co-reference set of any sub-document represent the non-co-reference scores of two entity mentions in the any sub-document with respect to the any co-reference set;

[0012] Calculate the global co-reference total score of any two entity mentions in the Internet document based on the co-reference total score and the non-co-reference total score of two entity mentions in each sub-document.

[0013] According to an Internet online common sense extraction method provided by the present invention, the calculating the global co-reference total score of any two entity mentions in the Internet document based on the co-reference total score and the non-co-reference total score of two entity mentions in each sub-document includes:

[0014] Obtain any two entity mentions belonging to the same sub-document, calculate the global co-reference total score of the any two entity mentions based on the co-reference total score and the non-co-reference total score of the any two entity mentions in each sub-document, and add the any two entity mentions to the mention pair set in the form of a mention pair;

[0015] For any first entity mention and second entity mention that do not belong to the same sub - document, based on the global co - reference total score of the mention pairs containing the first entity mention in the set of mention pairs, and the global co - reference total score of the mention pairs containing the second entity mention in the set of mention pairs, determine an intermediate mention, and based on the global co - reference total score between the first entity mention and the intermediate mention and the global co - reference total score between the intermediate mention and the second entity mention, determine the global co - reference total score of the first entity mention and the second entity mention; wherein, the intermediate mention satisfies the condition that the product of the global co - reference total score between the first entity mention and the intermediate mention and the global co - reference total score between the intermediate mention and the second entity mention is the largest.

[0016] According to an Internet online common sense extraction method provided by the present invention, the determining the global co - reference set based on the global co - reference total score of any two entity mentions in the Internet document includes:

[0017] Based on the global co - reference total score of any two entity mentions in the Internet document, determine a judgment threshold;

[0018] Based on the global co - reference total score of any two entity mentions in the Internet document and the judgment threshold, construct an entity mention graph; wherein, the nodes in the entity mention graph correspond to any entity mention, and if the global co - reference total score of any two entity mentions is higher than the judgment threshold, there is a connection edge between the nodes corresponding to the any two entity mentions;

[0019] Based on a graph convolutional neural network, obtain the node vector representation of each node in the entity mention graph, and based on the similarity between the node vector representations of each node in the entity mention graph, determine the global co - reference set.

[0020] According to an Internet online common sense extraction method provided by the present invention, the performing entity relationship extraction on the updated document includes:

[0021] Append the relationship names of each preset relationship to the end of the updated document to obtain a network input text;

[0022] Based on an encoding network, encode the network input text to obtain the entity representations of each entity and the relationship representations of each preset relationship in the network input text; the encoding network is constructed based on the Bert model;

[0023] Based on the relationship representations of each preset relationship, construct a relationship representation matrix;

[0024] For any head entity and any tail entity, determine the attention scores of the any head entity and the attention scores of the any tail entity calculated by each self-attention module in the last Transformer layer of the encoding network;

[0025] Based on the attention scores of the any head entity and the attention scores of the any tail entity calculated by each self-attention module, determine the entity attention score;

[0026] Determine the product of the entity attention score and the relation representation matrix as the relation attention representation matrix corresponding to the any head entity and the any tail entity;

[0027] Based on the relation attention representation matrix corresponding to the any head entity and the any tail entity and the entity representations of the any head entity and the any tail entity, use a classifier to determine the relation type between the any head entity and the any tail entity.

[0028] According to an Internet online common sense extraction method provided by the present invention, the using a classifier to determine the relation type between the any head entity and the any tail entity based on the relation attention representation matrix corresponding to the any head entity and the any tail entity and the entity representations of the any head entity and the any tail entity includes:

[0029] Fuse the relation attention representation matrix corresponding to the any head entity and the any tail entity with the entity representation of the any head entity and the entity representation of the any tail entity respectively to obtain the relation fusion entity representation of the any head entity and the relation fusion entity representation of the any tail entity;

[0030] Based on the classifier, perform relation classification on the relation fusion entity representation of the any head entity and the relation fusion entity representation of the any tail entity to obtain the relation type between the any head entity and the any tail entity.

[0031] According to an Internet online common sense extraction method provided by the present invention, the determining the entity attention score based on the attention scores of the any head entity and the attention scores of the any tail entity calculated by each self-attention module includes:

[0032] Multiply the attention scores of the any head entity and the attention scores of the any tail entity calculated by any self-attention module to obtain the single attention score corresponding to the any self-attention module;

[0033] Add the single attention scores corresponding to each self-attention module to obtain the entity attention score.

[0034] The present invention also provides an Internet online common sense extraction device, including:

[0035] A segmentation unit for segmenting Internet documents to obtain a plurality of sub-documents;

[0036] A local coreference resolution unit for performing coreference resolution and entity recognition on any sub-document to obtain a plurality of coreference sets and the entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of the any sub-document; wherein, the elements in the coreference sparse matrix corresponding to any coreference set represent the coreference scores of two entity mentions in the any sub-document with respect to the any coreference set;

[0037] A global coreference resolution unit for calculating the total coreference score of two entity mentions in the any sub-document based on the coreference sparse matrix corresponding to each coreference set of the any sub-document, calculating the global total coreference score of any two entity mentions in the Internet document based on the total coreference scores of two entity mentions in each sub-document, and determining a global coreference set based on the global total coreference score of any two entity mentions in the Internet document;

[0038] A knowledge extraction unit for replacing the entity mentions in the corresponding global coreference set based on the entity names of the entities in each global coreference set to obtain an updated document, extracting entity relationships from the updated document to obtain triples of the Internet document, and updating a knowledge base based on the triples of the Internet document.

[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the Internet online common sense extraction method as described in any one of the above.

[0040] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the Internet online common sense extraction method as described in any one of the above.

[0041] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the Internet online common sense extraction method as described in any one of the above.

[0042] The Internet online common sense extraction method, device, electronic device and storage medium provided by the present invention perform coreference resolution and entity recognition on sub-documents of Internet documents to obtain multiple coreference sets and entities existing in each coreference set, construct a coreference sparse matrix corresponding to each coreference set of each sub-document, and then calculate the coreference total score of two entity mentions in the sub-document based on the coreference sparse matrix corresponding to each coreference set of any sub-document. Based on the coreference total scores of two entity mentions in each sub-document, calculate the global coreference total score of any two entity mentions in the Internet document, and determine the global coreference set based on the global coreference total score of any two entity mentions in the Internet document. Thus, replace the entity mentions in the corresponding global coreference set with the entity names of the entities in each global coreference set to obtain an updated document, perform entity relationship extraction on the updated document to obtain triples of the Internet document, and update the knowledge base based on the triples of the Internet document, overcoming the long-distance dependence problem existing in long texts, improving the entity relationship extraction ability of long texts, and enhancing the knowledge extraction accuracy for Internet documents. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 is a flowchart of the Internet online common sense extraction method provided by the present invention;

[0045] Figure 2 is a flowchart of the global coreference total score calculation method provided by the present invention;

[0046] Figure 3 is a flowchart of the global coreference set construction method provided by the present invention;

[0047] Figure 4 is a structural schematic diagram of the Internet online common sense extraction device provided by the present invention;

[0048] Figure 5 is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0050] Figure 1 is a schematic flowchart of the Internet online common sense extraction method provided by the present invention. As Figure 1 shown, the method includes:

[0051] Step 110: Segment the Internet document to obtain multiple sub-documents;

[0052] Step 120: Perform coreference resolution and entity recognition on any one of the sub-documents to obtain multiple coreference sets and the entities existing in each coreference set, and construct a coreference sparse matrix corresponding to each coreference set of the any one of the sub-documents; wherein, the elements in the coreference sparse matrix corresponding to any one coreference set represent the coreference scores of two entity mentions in the any one of the sub-documents with respect to the any one coreference set.

[0053] Step 130: Based on the coreference sparse matrix corresponding to each coreference set of any one of the sub-documents, calculate the total coreference score of two entity mentions in the any one of the sub-documents. Based on the total coreference scores of two entity mentions in each sub-document, calculate the global total coreference score of any two entity mentions in the Internet document, and determine the global coreference set based on the global total coreference score of any two entity mentions in the Internet document.

[0054] Step 140: Based on the entity names of the entities in each global coreference set, replace the entity mentions in the corresponding global coreference set to obtain an updated document, perform entity relationship extraction on the updated document to obtain the triples of the Internet document, and update the knowledge base based on the triples of the Internet document.

[0055] Specifically, for a relatively long Internet document, it can be segmented into multiple shorter sub-documents. For each sub-document, coreference resolution can be performed on each sub-document based on a coreference resolution model (such as the Maverick model) to obtain multiple coreference sets respectively corresponding to each sub-document. Among them, any coreference set corresponding to a sub-document contains entity mentions that refer to the same entity object in that sub-document. In addition, entity recognition can be performed on each sub-document based on an entity recognition model, so as to identify the entities in the coreference sets of each sub-document. For any sub-document, a coreference sparse matrix corresponding to each coreference set of that sub-document can be constructed. Here, any element in the coreference sparse matrix corresponding to any coreference set of a sub-document represents the coreference score of two entity mentions in that sub-document relative to that coreference set. In some embodiments, if both of these two entity mentions are in that coreference set, the coreference score of these two entity mentions relative to that coreference set is 1, otherwise it is 0.

[0056] Based on the coreference sparse matrices corresponding to the coreference sets of any sub-document, the coreference total score of two entity mentions in that sub-document can be calculated. Among them, the coreference total score of any two entity mentions in that sub-document indicates the probability that these two entity mentions refer to the same entity within the scope of that sub-document. In some embodiments, the coreference scores of these two entity mentions relative to each coreference set (which can be obtained from the coreference sparse matrices corresponding to each coreference set) can be accumulated to obtain the coreference total score of these two entity mentions.

[0057] Since the length of Internet documents is relatively long, the key information related to the same entity may span multiple paragraphs, and there will be various pronouns or synonymous phrases used to describe the same entity. In order to improve the accuracy of entity relationship extraction in long texts and solve the long-distance dependence problem in long texts, further coreference resolution processing can be performed globally on the long text. Specifically, based on the coreference total scores of two entity mentions in each sub-document, the global coreference total score of any two entity mentions in the Internet document can be calculated, and based on the global coreference total score of any two entity mentions in the Internet document, the global coreference set can be determined. Among them, the global coreference total score of any two entity mentions in the Internet document indicates the probability that these two entity mentions refer to the same entity within the entire scope of the Internet document.

[0058] In some embodiments, in order to further improve the accuracy of coreference resolution, a non-coreference sparse matrix corresponding to each coreference set of each sub-document can be constructed. Among them, the elements in the non-coreference sparse matrix corresponding to any coreference set of any sub-document represent the non-coreference scores of two entity mentions in the sub-document with respect to the coreference set. In some embodiments, if at least one of the two entity mentions is not in the coreference set, the non-coreference score of the two entity mentions with respect to the coreference set is 1, otherwise it is 0. Based on the non-coreference sparse matrices corresponding to each coreference set of each sub-document, the total non-coreference score of two entity mentions in each sub-document can be calculated. Among them, the total non-coreference score of any two entity mentions in the sub-document indicates the probability that the two entity mentions refer to different entities within the scope of the sub-document. In some embodiments, the non-coreference scores of the two entity mentions with respect to each coreference set (which can be obtained from the non-coreference sparse matrix corresponding to each coreference set) can be accumulated to obtain the total non-coreference score of the two entity mentions. Subsequently, based on the coreference total score and the non-coreference total score of two entity mentions in each sub-document, the coreference total score and the non-coreference total score are used for complementarity, so as to accurately calculate the global coreference total score of any two entity mentions in the Internet document.

[0059] In some other embodiments, as Figure 2 shown, the following method can be used to calculate the global coreference total score of any two entity mentions in the Internet document:

[0060] Step 210, obtain any two entity mentions belonging to the same sub-document, calculate the global coreference total score of the any two entity mentions based on the coreference total score and the non-coreference total score of the any two entity mentions in each sub-document, and add the any two entity mentions to the mention pair set in the form of a mention pair;

[0061] Step 220, for a first entity mention and a second entity mention that do not belong to the same sub-document, based on the global coreference total score of the mention pair containing the first entity mention in the mention pair set, and the global coreference total score of the mention pair containing the second entity mention in the mention pair set, determine an intermediate mention, and based on the global coreference total score of the first entity mention and the intermediate mention and the global coreference total score of the intermediate mention and the second entity mention, determine the global coreference total score of the first entity mention and the second entity mention; among them, the intermediate mention satisfies the condition that the product of the global coreference total score of the first entity mention and the intermediate mention and the global coreference total score of the intermediate mention and the second entity mention is the largest.

[0062] Specifically, entity mention pairs <mentioni, mentionj> belonging to the same sub-document (which can be any one or more sub-documents) can be screened out first. For any entity mention pair belonging to the same sub-document, the total coreference scores (denoted as co1, co2,..., con, where n is the number of sub-documents) and non-coreference scores (denoted as nco1, nco2,..., ncon) of these two entity mentions in each sub-document can be determined. It should be noted that for any sub-document, if these two entity mentions do not appear in the sub-document simultaneously, the total coreference score of these two entity mentions in the sub-document is 0, and the non-coreference score is equal to the number of coreference sets in the sub-document. Subsequently, based on the total coreference scores and non-coreference scores of these two entity mentions in each sub-document, the global coreference score of these two entity mentions can be calculated, and these two entity mentions are added to the mention pair set in the form of a mention pair (i.e., <mentioni, mentionj>). Among them, the sum of the total coreference scores of these two entity mentions in each sub-document can be determined as the first parameter, the sum of the sum of the total coreference scores of these two entity mentions in each sub-document and the sum of the non-coreference scores of these two entity mentions in each sub-document can be determined as the second parameter, and the ratio of the first parameter to the second parameter is determined as the global coreference score of these two entity mentions.

[0063] After calculating the global coreference scores of all entity mention pairs belonging to the same sub-document in the above manner, for any first entity mention and second entity mention that do not belong to the same sub-document, the global coreference scores of the first entity mention and the second entity mention can be indirectly determined based on the global coreference scores of each entity mention pair belonging to the same sub-document calculated in the previous step. Specifically, an intermediate mention can be determined based on the global coreference scores of the mention pairs containing the first entity mention in the above mention pair set and the global coreference scores of the mention pairs containing the second entity mention in the above mention pair set. Among them, the intermediate mention satisfies the condition that the product of the global coreference score between the first entity mention and the intermediate mention and the global coreference score between the intermediate mention and the second entity mention is the largest. It can be seen that the mention pair composed of the first entity mention and the intermediate mention and the mention pair composed of the second entity mention and the intermediate mention are both in the above mention pair set. That is, the intermediate mention satisfies the condition: argmax mentionp(wscore(mention1, mentionp) × wscore(mentionp, mention2)), where mention1 and mention2 are the first entity mention and the second entity mention, mentionp is another entity mention, and wscore() is the global coreference total score. Subsequently, determine the product of the global coreference total score between the first entity mention and the intermediate mention and the global coreference total score between the intermediate mention and the second entity mention as the global coreference total score between the first entity mention and the second entity mention.

[0064] After obtaining the global coreference total score between any two entity mentions in the Internet document, the global coreference set can be determined accordingly. The entity mentions in any global coreference set refer to the same entity within the entire scope of the Internet document. In some embodiments, two entity mentions with a global coreference total score exceeding a certain threshold can be placed in the same global coreference set. Considering that the calculation of the global coreference total score between any two entity mentions in the Internet document is based on the understanding of each sub-document (i.e., short text) by the existing coreference resolution model, for some complex long texts (especially long texts with implicit inferences), its global understanding ability for such complex long texts may still be slightly lacking. Therefore, in order to further improve the analysis ability of long texts, in some other embodiments, graph theory can be used for coreference resolution. Specifically, as Figure 3 shown, the global coreference set can be determined in the following manner:

[0065] Step 310, determine a judgment threshold based on the global coreference total score between any two entity mentions in the Internet document;

[0066] Step 320, construct an entity mention graph based on the global coreference total score between any two entity mentions in the Internet document and the judgment threshold; wherein, the nodes in the entity mention graph correspond to any entity mention, and if the global coreference total score between any two entity mentions is higher than the judgment threshold, there is a connection edge between the nodes corresponding to the any two entity mentions;

[0067] Step 330, obtain the node vector representation of each node in the entity mention graph based on a graph convolutional neural network, and determine the global coreference set based on the similarity between the node vector representations of each node in the entity mention graph.

[0068] Here, the average value of the global co-reference total scores of any two entity mentions in the Internet document can be determined as the judgment threshold. When constructing the entity mention graph, all entity mentions are used as nodes in the entity mention graph. Further, based on the global co-reference total scores of any two entity mentions in the Internet document and the above-mentioned judgment threshold, connections are established for each entity mention. Specifically, if the global co-reference total score of any two entity mentions is higher than the above-mentioned judgment threshold, a connection edge is established between the nodes corresponding to these two entity mentions. Subsequently, based on the graph convolutional neural network, the features of each node itself and the structural relationship between nodes in the entity mention graph are automatically learned, so as to obtain the node vector representation of each node in the entity mention graph. In some embodiments, for the node corresponding to any entity mention, the encoding features of the entity mention and its context can be obtained based on a pre-trained language model (such as Bert), and the encoding features of the entity mention and its context are fused to obtain the node features of the node, so as to improve the accuracy of node semantic extraction. By inputting the node features of each node and the adjacency matrix corresponding to each node into the graph convolutional neural network, the node vector representation of each node output by the graph convolutional neural network can be obtained. Based on the similarity between the node vector representations of each node in the entity mention graph, the entity mentions corresponding to the nodes with similarity higher than the preset similarity threshold can be placed in the same global co-reference set, so as to determine the global co-reference set.

[0069] After obtaining the global co-reference set, the relationships between entities can be further extracted. In order to overcome the problem of long-distance dependence in relation extraction, in the Internet document, based on the entity names of the entities in any global co-reference set, other entity mentions belonging to the same global co-reference set can be replaced to obtain an updated document, and this updated document is used as the basis for entity relation extraction. Among them, the entities in the global co-reference set can be determined according to the previous entity recognition results. When performing entity relation extraction on the updated document, the relation names of each preset relation can be spliced to the end of the updated document to obtain the network input text. Subsequently, based on the pre-trained encoding network, the network input text is encoded to obtain the entity representations of each entity and the relation representations of each preset relation in the network input text. Among them, the encoding network is constructed based on the Bert model. After inputting the network input text into the encoding network, the encoding network will output the vector representations corresponding to each token in the network input text. For any entity, according to the position of the entity in the network input text (since the entity mentions pointing to the same entity are replaced above, there are multiple positions), all tokens corresponding to the entity in the network input text can be obtained, so as to fuse the vector representations of the corresponding tokens to obtain the entity representation of the entity. For any preset relation, the vector representation of the token corresponding to the preset relation in the network input text can be directly obtained and then fused to obtain the relation representation of the preset relation.

[0070] Based on the relationship representations of various preset relationships, a relationship representation matrix can be constructed. For any head entity and any tail entity, determine the attention scores of the head entity and the attention scores of the tail entity calculated by each self-attention module in the last Transformer layer of the encoding network. Among them, since the text input into the encoding network includes the updated document and various preset relationships, the attention scores of the head entity and the tail entity calculated by each self-attention module can indicate the degree of correlation between the head entity and the tail entity and each preset relationship. Based on the attention scores of the head entity and the attention scores of the tail entity calculated by each self-attention module, determine the entity attention score. In some embodiments, the attention scores of the head entity and the attention scores of the tail entity calculated by any self-attention module can be multiplied to obtain the single attention score corresponding to the self-attention module, and then the single attention scores corresponding to each self-attention module are added to obtain the entity attention score. Subsequently, determine the product of the entity attention score and the relationship representation matrix as the relationship attention representation matrix corresponding to the head entity and the tail entity to introduce the influence of the head entity and the tail entity into the relationship representation.

[0071] Based on the relationship attention representation matrix corresponding to the head entity and the tail entity and the entity representations of the head entity and the tail entity, a classifier can be used to determine the relationship type between the head entity and the tail entity by comprehensively considering the semantic information of the corresponding entities contained in the entity representations of the head entity and the tail entity and the semantic information of the relationships related to the head and tail entities contained in the relationship attention representation matrix. In some embodiments, the relationship attention representation matrix corresponding to the head entity and the tail entity can be respectively fused with the entity representation of the head entity and the entity representation of the tail entity to obtain the relationship fusion entity representation of the head entity and the relationship fusion entity representation of the tail entity, and then the classifier performs relationship classification on the relationship fusion entity representation of the head entity and the relationship fusion entity representation of the tail entity to obtain the relationship type between the head entity and the tail entity. Among them, the following formula can be used to implement the fusion of the relationship attention representation matrix with the entity representation of the head entity and the entity representation of the tail entity:

[0072] f s = tanh(W s h s +C s,o W r1 )

[0073] f o = tanh(W o h o +C s,o W r2 )

[0074] Among them, f s and f o are respectively the relation fusion entity representations of the head entity and the relation fusion entity representations of the tail entity, h s and h o are respectively the entity representations of the head entity and the entity representations of the tail entity, C s,o is the above-mentioned relation attention representation matrix, and W s , W r1 , W o and W r2 are learnable weight matrices.

[0075] Based on the head entity, the tail entity, and the relationship between the head entity and the tail entity, a triple can be constructed. Based on the triples in the Internet documents, the knowledge base can be updated to achieve the common sense extraction of the Internet documents.

[0076] In summary, the method provided by the embodiments of the present invention obtains multiple coreference sets and the entities existing in each coreference set by performing coreference resolution and entity recognition on the sub-documents of the Internet document, constructs a coreference sparse matrix corresponding to each coreference set of each sub-document, and then calculates the total coreference score of two entity mentions in the sub-document based on the coreference sparse matrix corresponding to each coreference set of any sub-document. Based on the total coreference scores of two entity mentions in each sub-document, calculates the global total coreference score of any two entity mentions in the Internet document, and determines the global coreference set based on the global total coreference score of any two entity mentions in the Internet document. Thus, replaces the entity mentions in the corresponding global coreference set with the entity names of the entities in each global coreference set to obtain an updated document, performs entity relationship extraction on the updated document to obtain the triples of the Internet document, and updates the knowledge base based on the triples of the Internet document, overcoming the long-distance dependence problem existing in long texts, improving the entity relationship extraction ability of long texts, and improving the knowledge extraction accuracy for Internet documents.

[0077] Next, the Internet online common sense extraction device provided by the present invention will be described. The Internet online common sense extraction device described below can be correspondingly referred to the Internet online common sense extraction method described above.

[0078] Based on any of the above embodiments, Figure 4 is a schematic structural diagram of the Internet online common sense extraction device provided by the present invention. As Figure 4 shown, the device includes:

[0079] A splitting unit 410, configured to split the Internet document to obtain multiple sub-documents;

[0080] The local coreference resolution unit 420 is used to perform coreference resolution and entity recognition on any sub-document, obtain multiple coreference sets and the entities existing in each coreference set, and construct a coreference sparse matrix corresponding to each coreference set of the any sub-document; wherein, the elements in the coreference sparse matrix corresponding to any coreference set represent the coreference scores of two entity mentions in the any sub-document with respect to the any coreference set.

[0081] The global coreference resolution unit 430 is used to calculate the total coreference score of two entity mentions in the any sub-document based on the coreference sparse matrix corresponding to each coreference set of the any sub-document, calculate the global total coreference score of any two entity mentions in the Internet document based on the total coreference scores of two entity mentions in each sub-document, and determine the global coreference set based on the global total coreference score of any two entity mentions in the Internet document.

[0082] The knowledge extraction unit 440 is used to replace the entity mentions in the corresponding global coreference set with the entity names of the entities in each global coreference set to obtain an updated document, perform entity relationship extraction on the updated document to obtain the triples of the Internet document, and update the knowledge base based on the triples of the Internet document.

[0083] The device provided by the embodiment of the present invention performs coreference resolution and entity recognition on the sub-documents of the Internet document, obtains multiple coreference sets and the entities existing in each coreference set, constructs a coreference sparse matrix corresponding to each coreference set of each sub-document, and then calculates the total coreference score of two entity mentions in the sub-document based on the coreference sparse matrix corresponding to each coreference set of the any sub-document, calculates the global total coreference score of any two entity mentions in the Internet document based on the total coreference scores of two entity mentions in each sub-document, and determines the global coreference set based on the global total coreference score of any two entity mentions in the Internet document, so as to replace the entity mentions in the corresponding global coreference set with the entity names of the entities in each global coreference set to obtain an updated document, perform entity relationship extraction on the updated document to obtain the triples of the Internet document, and update the knowledge base based on the triples of the Internet document, overcomes the long-distance dependence problem existing in long texts, improves the entity relationship extraction ability of long texts, and improves the knowledge extraction accuracy for Internet documents.

[0084] Based on any of the above embodiments, calculating the global total coreference score of any two entity mentions in the Internet document based on the total coreference scores of two entity mentions in each sub-document includes:

[0085] Construct non-coreferential sparse matrices corresponding to the coreferential sets of each sub-document, and calculate the total non-coreferential scores of two entity mentions in each sub-document based on the non-coreferential sparse matrices corresponding to the coreferential sets of each sub-document; wherein, the elements in the non-coreferential sparse matrix corresponding to any coreferential set of any sub-document represent the non-coreferential scores of two entity mentions in the any sub-document with respect to the any coreferential set.

[0086] Based on the total coreferential scores and total non-coreferential scores of two entity mentions in each sub-document, calculate the global coreferential total score of any two entity mentions in the Internet document.

[0087] Based on any of the above embodiments, the calculating the global coreferential total score of any two entity mentions in the Internet document based on the total coreferential scores and total non-coreferential scores of two entity mentions in each sub-document includes:

[0088] Obtain any two entity mentions belonging to the same sub-document, calculate the global coreferential total score of the any two entity mentions based on the total coreferential scores and total non-coreferential scores of the any two entity mentions in each sub-document, and add the any two entity mentions to the mention pair set in the form of a mention pair.

[0089] For a first entity mention and a second entity mention that do not belong to the same sub-document, based on the global coreferential total scores of the mention pairs containing the first entity mention in the mention pair set and the global coreferential total scores of the mention pairs containing the second entity mention in the mention pair set, determine an intermediate mention, and based on the global coreferential total score between the first entity mention and the intermediate mention and the global coreferential total score between the intermediate mention and the second entity mention, determine the global coreferential total score of the first entity mention and the second entity mention; wherein, the intermediate mention satisfies the condition that the product of the global coreferential total score between the first entity mention and the intermediate mention and the global coreferential total score between the intermediate mention and the second entity mention is the largest.

[0090] Based on any of the above embodiments, the determining the global coreferential set based on the global coreferential total score of any two entity mentions in the Internet document includes:

[0091] Determine a judgment threshold based on the global coreferential total score of any two entity mentions in the Internet document.

[0092] Construct an entity mention graph based on the global coreferential total score of any two entity mentions in the Internet document and the judgment threshold; wherein, the nodes in the entity mention graph correspond to any entity mention, and if the global coreferential total score of any two entity mentions is higher than the judgment threshold, there is a connection edge between the nodes corresponding to the any two entity mentions.

[0093] Obtain the node vector representations of each node in the entity mention graph based on a graph convolutional neural network, and determine the global coreference set based on the similarity between the node vector representations of each node in the entity mention graph.

[0094] Based on any of the above embodiments, the entity relationship extraction from the updated document includes:

[0095] Splice the relationship names of each preset relationship to the end of the updated document to obtain the network input text;

[0096] Encode the network input text based on an encoding network to obtain the entity representations of each entity and the relationship representations of each preset relationship in the network input text; the encoding network is constructed based on the Bert model;

[0097] Construct a relationship representation matrix based on the relationship representations of each preset relationship;

[0098] For any head entity and any tail entity, determine the attention scores of the any head entity and the any tail entity calculated by each self-attention module in the last Transformer layer of the encoding network;

[0099] Determine the entity attention score based on the attention scores of the any head entity and the any tail entity calculated by each self-attention module;

[0100] Determine the product of the entity attention score and the relationship representation matrix as the relationship attention representation matrix corresponding to the any head entity and the any tail entity;

[0101] Based on the relationship attention representation matrix corresponding to the any head entity and the any tail entity and the entity representations of the any head entity and the any tail entity, use a classifier to determine the relationship type between the any head entity and the any tail entity.

[0102] Based on any of the above embodiments, the determining the relationship type between the any head entity and the any tail entity by using a classifier based on the relationship attention representation matrix corresponding to the any head entity and the any tail entity and the entity representations of the any head entity and the any tail entity includes:

[0103] Fuse the relationship attention representation matrix corresponding to the any head entity and the any tail entity with the entity representation of the any head entity and the entity representation of the any tail entity respectively to obtain the relationship-fused entity representation of the any head entity and the relationship-fused entity representation of the any tail entity;

[0104] Performing relation classification on the relation fusion entity representation of any head entity and the relation fusion entity representation of any tail entity based on the classifier to obtain the relation type between the any head entity and the any tail entity.

[0105] Based on any of the above embodiments, determining an entity attention score based on the attention scores of any head entity and any tail entity calculated by each self-attention module, including:

[0106] Multiplying the attention score of any head entity and the attention score of any tail entity calculated by any self-attention module to obtain a single attention score corresponding to the any self-attention module;

[0107] Adding the single attention scores corresponding to each self-attention module to obtain the entity attention score.

[0108] Figure 5 It is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a memory 520, a communication interface 530, and a communication bus 540. Among them, the processor 510, the memory 520, and the communication interface 530 complete mutual communication through the communication bus 540. The processor 510 may call logical instructions in the memory 520 to execute an Internet online common sense extraction method, which includes: segmenting Internet documents to obtain a plurality of sub-documents; performing coreference resolution and entity recognition on any sub-document to obtain a plurality of coreference sets and entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of the any sub-document; where the elements in the coreference sparse matrix corresponding to any coreference set represent the coreference scores of two entity mentions in the any sub-document with respect to the any coreference set; based on the coreference sparse matrices corresponding to each coreference set of any sub-document, calculating the total coreference score of two entity mentions in the any sub-document, based on the total coreference scores of two entity mentions in each sub-document, calculating the global total coreference score of any two entity mentions in the Internet document, and based on the global total coreference score of any two entity mentions in the Internet document, determining a global coreference set; replacing entity mentions in the corresponding global coreference set with the entity names of entities in each global coreference set to obtain an updated document, and performing entity relation extraction on the updated document to obtain triples of the Internet document, and updating the knowledge base based on the triples of the Internet document.

[0109] In addition, when the logical instructions in the above-mentioned memory 520 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0110] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the Internet online common sense extraction method provided by the above-mentioned various methods. The method includes: splitting an Internet document to obtain a plurality of sub-documents; performing coreference resolution and entity recognition on any one of the sub-documents to obtain a plurality of coreference sets and the entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of the any one of the sub-documents; wherein, the elements in the coreference sparse matrix corresponding to any one coreference set represent the coreference scores of two entity mentions in the any one of the sub-documents with respect to the any one coreference set; based on the coreference sparse matrices corresponding to each coreference set of any one of the sub-documents, calculating the total coreference score of two entity mentions in the any one of the sub-documents, based on the total coreference scores of two entity mentions in each sub-document, calculating the global total coreference score of any two entity mentions in the Internet document, and based on the global total coreference score of any two entity mentions in the Internet document, determining the global coreference set; based on the entity names of the entities in each global coreference set, replacing the entity mentions in the corresponding global coreference set to obtain an updated document, and performing entity relationship extraction on the updated document to obtain the triples of the Internet document, and updating the knowledge base based on the triples of the Internet document.

[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the Internet online common sense extraction method provided above. The method includes: splitting an Internet document to obtain a plurality of sub-documents; performing coreference resolution and entity recognition on any one of the sub-documents to obtain a plurality of coreference sets and the entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of the any one of the sub-documents; wherein, the elements in the coreference sparse matrix corresponding to any one coreference set represent the coreference scores of two entity mentions in the any one of the sub-documents with respect to the any one coreference set; based on the coreference sparse matrices corresponding to each coreference set of any one of the sub-documents, calculating the total coreference score of two entity mentions in the any one of the sub-documents, based on the total coreference scores of two entity mentions in each sub-document, calculating the global total coreference score of any two entity mentions in the Internet document, and based on the global total coreference score of any two entity mentions in the Internet document, determining the global coreference set; replacing the entity mentions in the corresponding global coreference set with the entity names of the entities in each global coreference set to obtain an updated document, and performing entity relationship extraction on the updated document to obtain the triples of the Internet document, and updating the knowledge base based on the triples of the Internet document.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for extracting online common sense from the Internet, characterized in that: include: Segment the Internet document to obtain multiple sub-documents; Perform coreference resolution and entity recognition on any sub-document to obtain multiple coreference sets and entities in each coreference set, and construct a coreference sparse matrix corresponding to each coreference set of any sub-document; wherein the elements in the coreference sparse matrix corresponding to any coreference set represent the coreference scores of two entity mentions in any sub-document relative to any coreference set; Based on the coreference sparse matrix corresponding to each coreference set of any sub-document, calculating the coreference total score of two entity mentions in any sub-document, calculating the global coreference total score of any two entity mentions in the Internet document based on the coreference total score of the two entity mentions in each sub-document, and determining the global coreference set based on the global coreference total score of any two entity mentions in the Internet document; Based on the entity names of the entities in each global coreference set, entity mentions in the corresponding global coreference set are replaced to obtain an updated document, and entity relations are extracted from the updated document to obtain triples of the Internet document, and a knowledge base is updated based on the triples of the Internet document.

2. The method for extracting online common sense from the Internet according to claim 1, characterized in that: The calculating the global coreference score of any two entity mentions in the Internet document based on the coreference score of the two entity mentions in each sub-document includes: Constructing a non-coreference sparse matrix corresponding to each coreference set of each sub-document, and calculating the total non-coreference score of the two entity mentions in each sub-document based on the non-coreference sparse matrix corresponding to each coreference set of each sub-document; wherein the elements in the non-coreference sparse matrix corresponding to any coreference set of any sub-document represent the non-coreference score of the two entity mentions in any sub-document relative to the any coreference set; Based on the total coreference scores and the total non-coreference scores of the two entity mentions in each sub-document, a global coreference score of any two entity mentions in the Internet document is calculated.

3. The method for extracting online common sense from the Internet according to claim 2, characterized in that: The calculating of the global coreference score of any two entity mentions in the Internet document based on the coreference score and the non-coreference score of the two entity mentions in each sub-document comprises: Obtain any two entity mentions belonging to the same sub-document, calculate the global coreference score of the any two entity mentions based on the coreference score and the non-coreference score of the any two entity mentions in each sub-document, and add the any two entity mentions to the mention pair set in the form of a mention pair; For any first entity mention and second entity mention that do not belong to the same sub-document, the intermediate mention is determined based on the global coreference score of the mention pairs in the mention pair set that includes the first entity mention, and the global coreference score of the mention pairs in the mention pair set that includes the second entity mention, and the global coreference score of the first entity mention and the second entity mention is determined based on the global coreference score of the first entity mention and the intermediate mention and the global coreference score of the intermediate mention and the second entity mention; wherein the intermediate mention satisfies the condition that the product of the global coreference score of the first entity mention and the intermediate mention and the global coreference score of the intermediate mention and the second entity mention is the largest.

4. The method for extracting online common sense from the Internet according to claim 1, characterized in that: The determining of a global coreference set based on the total global coreference score of any two entity mentions in the Internet document comprises: Determining a judgment threshold based on the total global coreference score of any two entity mentions in the Internet document; Based on the total global coreference score of any two entity mentions in the Internet document and the judgment threshold, constructing an entity mention graph; wherein a node in the entity mention graph corresponds to any entity mention, and if the total global coreference score of any two entity mentions is higher than the judgment threshold, then there is a connecting edge between the nodes corresponding to the any two entity mentions; A node vector representation of each node in the entity mention graph is obtained based on a graph convolutional neural network, and the global coreference set is determined based on the similarity between the node vector representations of each node in the entity mention graph.

5. The method for extracting online common sense from the Internet according to any one of claims 1 to 4, characterized in that: The extracting entity relationships from the update document includes: splicing the relationship names of the preset relationships to the end of the update document to obtain a network input text; Encoding the network input text based on a coding network to obtain entity representations of each entity in the network input text and relationship representations of each preset relationship; the coding network is constructed based on a Bert model; Constructing a relational representation matrix based on the relational representations of each preset relation; For any head entity and any tail entity, determine the attention score of any head entity and the attention score of any tail entity calculated by each self-attention module of the last Transformer layer in the encoding network; Determine an entity attention score based on the attention score of any head entity and the attention score of any tail entity calculated by each self-attention module; Determine the product of the entity attention score and the relationship representation matrix as the relationship attention representation matrix corresponding to any head entity and any tail entity; Based on the relationship attention representation matrix corresponding to any head entity and any tail entity and the entity representation of any head entity and any tail entity, a classifier is used to determine the relationship type between any head entity and any tail entity.

6. The method for extracting online common sense from the Internet according to claim 5, characterized in that: The determining the relationship type between any head entity and any tail entity using a classifier based on the relationship attention representation matrix corresponding to any head entity and any tail entity and the entity representation of any head entity and any tail entity includes: The relation attention representation matrices corresponding to any head entity and any tail entity are respectively fused with the entity representation of any head entity and the entity representation of any tail entity to obtain the relation fused entity representation of any head entity and the relation fused entity representation of any tail entity; Based on the classifier, relationship classification is performed on the relationship fusion entity representation of any head entity and the relationship fusion entity representation of any tail entity to obtain the relationship type between any head entity and any tail entity.

7. The method for extracting online common sense from the Internet according to claim 5, characterized in that: Determining an entity attention score based on the attention score of any head entity and the attention score of any tail entity calculated by each self-attention module includes: Multiply the attention score of any head entity and the attention score of any tail entity calculated by any self-attention module to obtain a single attention score corresponding to any self-attention module; The single attention scores corresponding to each self-attention module are added together to obtain the entity attention score.

8. An Internet online common sense extraction device, characterized in that: include: A segmentation unit, used for segmenting an Internet document to obtain a plurality of sub-documents; A local coreference resolution unit, used for performing coreference resolution and entity recognition on any sub-document, obtaining multiple coreference sets and entities existing in each coreference set, and constructing a coreference sparse matrix corresponding to each coreference set of any sub-document; wherein the elements in the coreference sparse matrix corresponding to any coreference set represent the coreference scores of two entity mentions in any sub-document relative to any coreference set; A global coreference resolution unit, configured to calculate a total coreference score of two entity mentions in any sub-document based on a coreference sparse matrix corresponding to each coreference set of any sub-document, calculate a total global coreference score of any two entity mentions in the Internet document based on the total coreference scores of the two entity mentions in each sub-document, and determine a global coreference set based on the total global coreference scores of any two entity mentions in the Internet document; A knowledge extraction unit is used to replace entity mentions in corresponding global co-reference sets based on entity names of entities in each global co-reference set to obtain updated documents, extract entity relationships from the updated documents to obtain triples of the Internet documents, and update the knowledge base based on the triples of the Internet documents.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for extracting online common sense from the Internet as claimed in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for extracting online common sense from the Internet as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Short text microblog co-reference resolution model for network feedback information monitoring

    CN118709687A

  • Entity-oriented multi-document abstract generation system and method

    CN119415684A