Relation Extraction Method, Device, Equipment and Storage Medium

By annotating and encoding the entities and keywords in the target text, using the self-attention mechanism to extract vectors, and combining with the classification network to determine the relationship, the problem of Bootstrapping's poor effect on complex contextual texts is solved, and a higher accuracy and generalization relationship extraction is achieved.

CN114281938BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111194402.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-06
Filing Date
2021-10-13
Publication Date
2025-07-11
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

When using Bootstrapping for relationship extraction in the prior art, the generalization of the template is insufficient, resulting in poor results on texts with complex contexts and the inability to accurately identify the relationship between entities.

Method used

By annotating entities and keywords in the target text, the encoding network and self-attention mechanism are used to extract the encoding representation vector of entities and the entity keyword representation vector, and the relationship between entities is determined in combination with the classification network to reduce the influence of interfering words in the text.

Benefits of technology

It improves the accuracy and generalization of relationship extraction, can provide more perfect relationship extraction results on complex context texts, reduces template dependence, and is suitable for most texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114281938B_ABST
    Figure CN114281938B_ABST
Patent Text Reader

Abstract

The present application discloses a relation extraction method, apparatus, device, and storage medium, which relate to technical fields such as artificial intelligence and intelligent transportation. The method includes: obtaining a target text containing a first entity and a second entity; annotating the first entity, the second entity, and keywords in the target text to obtain an annotated target text; performing encoding processing on the annotated target text to obtain an encoded representation vector corresponding to the first entity and an entity keyword representation vector, and an encoded representation vector corresponding to the second entity and an entity keyword representation vector; determining the relationship between the first entity and the second entity according to the encoded representation vector corresponding to the first entity and the entity keyword representation vector, and the encoded representation vector corresponding to the second entity and the entity keyword representation vector. The present application provides a relation extraction solution with stronger generalization ability, thereby helping to improve the perfection and accuracy of the relation extraction result.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202111037638.9 and the invention title "Relationship Extraction Method, Device, Equipment and Storage Medium" filed on September 6, 2021, the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the technical fields of artificial intelligence, intelligent transportation, etc., and particularly relate to a relationship extraction method, device, equipment and storage medium. Background Art

[0003] Relationship extraction refers to obtaining the relationships between various entities included in a text. For example, taking person relationship extraction as an example, two person names are identified from the text, and then based on the information included in the text, the person relationship between these two person names is determined, such as the relationship of children, spouse, etc.

[0004] In the related art, a solution for relationship extraction by constructing a semantic template using Bootstrapping is provided.

[0005] However, the main disadvantage of using Bootstrapping for relationship extraction is the insufficient generalization of the template, which has poor performance on texts with complex contexts and cannot obtain very perfect and correct relationships between entities. Summary of the Invention

[0006] Embodiments of this application provide a relationship extraction method, device, equipment and storage medium, providing a relationship extraction solution with stronger generalization, thereby helping to improve the perfection and accuracy of relationship extraction results. The technical solution is as follows:

[0007] According to one aspect of the embodiments of this application, a relationship extraction method is provided, and the method includes:

[0008] Obtain a target text containing a first entity and a second entity;

[0009] Annotate the first entity, the second entity and keywords in the target text to obtain an annotated target text; wherein, the keywords refer to the words in the target text that can reflect the relationship between the first entity and the second entity;

[0010] Perform encoding processing on the annotated target text to obtain an encoding representation vector and an entity keyword representation vector corresponding to the first entity, and an encoding representation vector and an entity keyword representation vector corresponding to the second entity; wherein, the encoding representation vector is used to reflect the feature information of the entity, and the entity keyword representation vector is used to reflect the correlation degree between the entity and the keyword;

[0011] Determine the relationship between the first entity and the second entity according to the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, and the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors.

[0012] According to one aspect of the embodiments of the present application, a method for training a relationship extraction model is provided. The method includes:

[0013] Obtain training samples of the relationship extraction model. The training samples include: sample texts containing a first entity and a second entity, and the true relationship between the first entity and the second entity;

[0014] Annotate the first entity, the second entity, and keywords in the sample text to obtain an annotated sample text; wherein, the keywords refer to the words in the sample text that can reflect the relationship between the first entity and the second entity;

[0015] Perform encoding processing on the annotated sample text through the encoding network of the relationship extraction model to obtain the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, and the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors; wherein, the encoded representation vectors are used to reflect the feature information of the entity, and the entity keyword representation vectors are used to reflect the correlation degree between the entity and the keyword;

[0016] Determine the predicted relationship between the first entity and the second entity through the classification network of the relationship extraction model according to the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, and the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors;

[0017] Determine the training loss of the relationship extraction model according to the true relationship and the predicted relationship, and adjust the network parameters of the relationship extraction model based on the training loss.

[0018] According to one aspect of the embodiments of the present application, a device for relationship extraction is provided. The device includes:

[0019] A text acquisition module, configured to acquire a target text containing a first entity and a second entity;

[0020] An annotation module, configured to annotate the first entity, the second entity, and keywords in the target text to obtain an annotated target text; wherein, the keywords refer to the words in the target text that can reflect the relationship between the first entity and the second entity;

[0021] An encoding processing module, configured to perform encoding processing on the labeled target text to obtain an encoded representation vector corresponding to the first entity and an entity keyword representation vector, and an encoded representation vector corresponding to the second entity and an entity keyword representation vector; wherein, the encoded representation vector is used to reflect the feature information of the entity, and the entity keyword representation vector is used to reflect the association degree between the entity and the keyword;

[0022] A relationship determination module, configured to determine the relationship between the first entity and the second entity according to the encoded representation vector corresponding to the first entity and the entity keyword representation vector, and the encoded representation vector corresponding to the second entity and the entity keyword representation vector.

[0023] According to one aspect of the embodiments of the present application, there is provided a training device for a relationship extraction model, the device includes:

[0024] A sample acquisition module, configured to acquire training samples of the relationship extraction model, the training samples including: a sample text containing a first entity and a second entity, and the true relationship between the first entity and the second entity;

[0025] A sample annotation module, configured to annotate the first entity, the second entity and keywords in the sample text to obtain a labeled sample text; wherein, the keyword refers to a word or phrase in the sample text that can reflect the relationship between the first entity and the second entity;

[0026] An encoding processing module, configured to perform encoding processing on the labeled sample text through the encoding network of the relationship extraction model to obtain an encoded representation vector corresponding to the first entity and an entity keyword representation vector, and an encoded representation vector corresponding to the second entity and an entity keyword representation vector; wherein, the encoded representation vector is used to reflect the feature information of the entity, and the entity keyword representation vector is used to reflect the association degree between the entity and the keyword;

[0027] A relationship determination module, configured to determine a predicted relationship between the first entity and the second entity through the classification network of the relationship extraction model according to the encoded representation vector corresponding to the first entity and the entity keyword representation vector, and the encoded representation vector corresponding to the second entity and the entity keyword representation vector;

[0028] A parameter adjustment module, configured to determine the training loss of the relationship extraction model according to the true relationship and the predicted relationship, and adjust the network parameters of the relationship extraction model based on the training loss.

[0029] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned relation extraction method or the training method of the above-mentioned relation extraction model.

[0030] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned relation extraction method or the training method of the above-mentioned relation extraction model.

[0031] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned relation extraction method or the training method of the above-mentioned relation extraction model.

[0032] The technical solutions provided by the embodiments of the present application at least include the following beneficial effects:

[0033] By annotating the entities and keywords in the target text, after obtaining the annotated target text, encoding processing is performed on the annotated target text to obtain the encoded representation vectors corresponding to the entities and the entity keyword representation vectors, so as to obtain the feature information of the entities in the target text and the correlation degree between the entities and the keywords. Then, based on the above information, the relationship between the entities is determined, thereby providing a relation extraction method that does not rely on relation templates, overcoming the problem of insufficient generalization of relation templates. The relation extraction method provided by the present application has strong generalization and can provide more perfect relation extraction results. Moreover, by extracting the entity keyword representation vectors and using the entity keyword representation vectors to characterize the correlation degree between the entities and the keywords, the influence of interfering words in the text on relation extraction is effectively reduced, thereby effectively solving the problem that Bootstrapping has poor effects in processing texts with complex contexts and improving the accuracy of relation extraction results.

[0034] At the same time, the relation extraction method provided by the present application only focuses on the entities and keywords in the target text and does not need to consider the positions of the entities and keywords in the target text, so it is applicable to most texts, effectively solving the deficiency of Bootstrapping in terms of versatility and enhancing the generalization of the relation extraction method. Brief Description of the Drawings

[0035] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1 is a schematic diagram of the solution implementation environment provided by an embodiment of the present application;

[0037] Figure 2 is a schematic diagram of the solution application scenario provided by an embodiment of the present application;

[0038] Figure 3 is a flowchart of the relationship extraction method provided by an embodiment of the present application;

[0039] Figure 4 is a schematic diagram of the text annotation provided by an embodiment of the present application;

[0040] Figure 5 is a flowchart of the relationship extraction method provided by another embodiment of the present application;

[0041] Figure 6 is a framework diagram of the relationship extraction model provided by an embodiment of the present application;

[0042] Figure 7 is a schematic diagram of obtaining the entity keyword representation vector based on the self-attention mechanism provided by an embodiment of the present application;

[0043] Figure 8 is a schematic diagram of the extraction of the relationship between people provided by an embodiment of the present application;

[0044] Figure 9 is a flowchart of the training method of the relationship extraction model provided by an embodiment of the present application;

[0045] Figure 10 is a framework diagram of the training of the relationship extraction model provided by an embodiment of the present application;

[0046] Figure 11 is a block diagram of the relationship extraction device provided by an embodiment of the present application;

[0047] Figure 12 is a block diagram of the training device of the relationship extraction model provided by an embodiment of the present application;

[0048] Figure 13 is a schematic diagram of the structure of the computer device provided by an embodiment of the present application. Detailed Implementation Modes

[0049] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe in detail the implementation manners of this application with reference to the accompanying drawings.

[0050] Artificial Intelligence (AI for short) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0051] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0052] Natural Language Processing (NLP for short) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life. So it has a close connection with the research of linguistics, but there are also important differences. Natural language processing does not generally study natural language, but aims to develop a computer system that can effectively achieve natural language communication, especially the software system therein. Thus, it is a part of computer science. Natural language processing is mainly applied in aspects such as machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, and speech recognition.

[0053] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0054] The technical solutions provided in the embodiments of this application involve technologies such as natural language processing and machine learning in artificial intelligence, and are specifically introduced and described through the following embodiments.

[0055] Please refer to Figure 1 , which shows a schematic diagram of the solution implementation environment provided in an embodiment of this application. This solution implementation environment can be implemented as a relation extraction system. This solution implementation environment may include a model training device 10 and a model using device 20.

[0056] The model training device 10 can be an electronic device such as a computer, a server, an intelligent robot, or some other electronic device with strong computing power. The model training device 10 is used to train the relation extraction model 30. In the embodiments of this application, the relation extraction model 30 is a model for extracting relations from text, and the model training device 10 can use machine learning methods to train this relation extraction model 30 to make it have better performance.

[0057] The above-mentioned trained relation extraction model 30 can be deployed in the model using device 20 for use to extract relations from text and obtain corresponding extraction results (i.e., the relations between entity pairs). The model using device 20 can be a terminal device such as a mobile phone, a computer, an intelligent TV, a multimedia playback device, a wearable device, a medical device, or a server. This application does not make any restrictions on this.

[0058] In some embodiments, as Figure 1 shown, the relation extraction model 30 includes an encoding network 31 and a classification network 32. Among them, the encoding network 31 is used to extract the encoding representation vector h(S) corresponding to the first entity and the entity keyword representation vector h kw (S), as well as the encoding representation vector h(O) corresponding to the second entity and the entity keyword representation vector h kw (O) from the target text containing the first entity and the second entity; the classification network 32 is used to perform classification prediction based on the above information extracted by the encoding network 31 to obtain the relation between the first entity and the second entity.

[0059] In the following method embodiments, the usage process and training process of the relationship extraction model 30 will be introduced in detail.

[0060] In some embodiments, take the application of the technical solution of this application to the extraction of character relationships in film and television dramas as an example. As Figure 2 shown, the application scenario of this solution may include: a terminal 210 and a server 220.

[0061] The terminal 210 may be an electronic device such as a mobile phone, a tablet computer, a PC (Personal Computer), a wearable device, a smart TV, a vehicle-mounted terminal, etc. A client for a target application for displaying character relationships may be installed and run in the terminal 210. For example, the target application may be a browser, a video application, an encyclopedia application, a search application, etc., and this application does not make any limitations thereto.

[0062] The server 220 may be a single server, a server cluster composed of multiple servers, or a cloud computing service center. The server 220 may be the background server of the above-mentioned target application, used to provide background services for the target application. Exemplarily, the above-mentioned relationship extraction model is deployed in the server 220. The server 220 performs relationship extraction on the film and television drama text through this relationship extraction model, and then sends the obtained extraction result (that is, the character relationship in the film and television drama) to the client of the target application for display.

[0063] The terminal 210 and the server 220 can communicate with each other through a network.

[0064] For example, the client of the target application may display an introduction interface 230 of a certain film and television drama character A. In this introduction interface 230, in addition to including the profile information 231 of the film and television drama character A, it may also include several character relationship cards 232. Each character relationship card 232 is used to display other film and television drama characters having a relevant relationship with the film and television drama character A, such as the lover, father-in-law, mother-in-law, etc. of the film and television drama character A. The above-mentioned character relationship card may also be referred to as a character KG (Knowledge Graph) card or other names, and this application does not make any limitations thereto. The character relationship card is a character information card containing character relationships obtained after sorting.

[0065] Of course, the technical solution provided by the embodiments of the present application can be applied not only to extracting the relationships among characters from film and television drama texts, but also to extracting the relationships among characters from other texts (such as novels, comics, web pages, etc.), or extracting other types of relationships, such as the relationships between characters and locations, between characters and events, between locations and events, etc. It can also be applied to scenarios such as supplementing the missing entity relationships in knowledge base question answering. The present application does not make any limitations in this regard.

[0066] The solution for relationship extraction using Bootstrapping provided by the related technology is as follows: Taking the extraction of character relationships as an example, given that Tom and Lucy are in a spousal relationship, texts such as "Tom's wife is Lucy" and "Lucy is Tom's wife" are extracted from a large number of texts based on this spousal relationship. Then, these texts are generalized to obtain relationship templates for spouses such as "XX's wife is XX" and "XX is XX's wife". Next, character relationships are extracted from a large number of texts using these relationship templates to obtain pairs of characters in a spousal relationship, such as "Jack and Lily", "Allen and Lara", etc. Then, more texts about spouses are extracted using these pairs of characters in a spousal relationship for generalization to obtain more relationship templates for spouses. These new templates are added to the old template library. After repeating this process multiple times, the template library is expanded. Based on this template library, character relationships are extracted, and a large number of pairs of characters in a spousal relationship can be obtained. However, the main drawback of using Bootstrapping for relationship extraction is the insufficient generalization of the templates, which results in poor performance on texts with complex contexts and cannot obtain very perfect and correct relationships between entities.

[0067] In the present application, after annotating the entities and keywords in the target text to obtain the annotated target text, the annotated target text is encoded to obtain the encoded representation vectors corresponding to the entities and the entity keyword representation vectors, thereby obtaining the feature information of the entities in the target text and the correlation degree between the entities and the keywords. Then, based on the above information, the relationships between the entities are determined, thus providing a relationship extraction method that does not rely on relationship templates, overcoming the problem of insufficient generalization of relationship templates. The relationship extraction method provided by the present application has strong generalization ability and can provide a more perfect relationship extraction result. Moreover, by extracting the entity keyword representation vectors and using these vectors to represent the correlation degree between the entities and the keywords, the influence of interfering words in the text on relationship extraction is effectively reduced, thus effectively solving the problem that Bootstrapping has poor performance on texts with complex contexts and improving the accuracy of the relationship extraction result.

[0068] Next, the technical solution of the present application will be introduced and illustrated through several embodiments.

[0069] Please refer to Figure 3 , which shows a flowchart of a relationship extraction method provided in an embodiment of the present application. The execution subject of each step of this method can be Figure 1 the model using device 20 in the solution implementation environment shown. This method may include the following steps (310-340):

[0070] Step 310, obtain a target text containing a first entity and a second entity.

[0071] An entity refers to an objectively existing and distinguishable thing. Generally speaking, entities are represented by nouns. For example, entities include but are not limited to personal names, place names, events, etc., and the present application does not make any limitations in this regard.

[0072] In some embodiments, the target text includes a first entity and a second entity for which the relationship is to be extracted. Among them, the first entity and the second entity are two different entities. In addition, the first entity and the second entity may belong to the same type. For example, both the first entity and the second entity are personal names; or, the first entity and the second entity may also belong to different types. For example, the first entity is a personal name and the second entity is an event.

[0073] Optionally, the first entity and the second entity are identified from the target text by using Named Entity Recognition (NER) technology. NER technology can identify entities with specific meanings in the text, mainly including personal names, place names, organization names, proper nouns, etc. NER is an important basic tool in application fields such as information extraction, question answering systems, syntactic analysis, and machine translation, and occupies an important position in the process of the practical application of natural language processing technology.

[0074] In some embodiments, the first entity and the second entity may be two entities specified in the text. In one example, assume that the target text is "Zhang San is Li Si's first love and Wang Wu's friend". First, use NER to identify the entities from the target text, such as including 3 entities: "Zhang San", "Li Si", and "Wang Wu". Assume that the relationship between "Zhang San" and "Li Si" is specified to be extracted. Then, the first entity is specified as "Zhang San" and the second entity is "Li Si", or the first entity can also be "Li Si" and the second entity is "Zhang San". "Wang Wu", as an entity that is not specified, will not become the first entity or the second entity.

[0075] In some embodiments, the first entity and the second entity can be any two entities not specified in the text. In one example, assume the target text is "Zhang San is Li Si's first love and Wang Wu's friend". First, use NER to identify the entities from the target text, such as including 3 entities: "Zhang San", "Li Si", and "Wang Wu". If the first entity and the second entity are not specified, then the first entity can be any one of "Zhang San", "Li Si", and "Wang Wu", and the second entity can also be any one of "Zhang San", "Li Si", and "Wang Wu", but the first entity and the second entity cannot be the same entity. For example, if the first entity is "Zhang San" and the second entity is "Li Si", extract the relationship between "Zhang San" and "Li Si" based on the content of the target text. Another example, if the first entity is "Zhang San" and the second entity is "Wang Wu", extract the relationship between "Zhang San" and "Wang Wu" based on the content of the target text.

[0076] The target text can be one or more sentences selected from the source text for extracting the relationship between entities. In some embodiments, taking the extraction of character relationships in a movie or TV drama as an example, the source text can be the text of a movie or TV drama, and the target text can be one or more sentences selected from the movie or TV drama text. Among them, the movie or TV drama text can be the script, lines, etc. of a certain movie or TV drama, and this application does not limit this.

[0077] It should be noted that generally speaking, when extracting entity relationships, two entities need to form an entity pair. For this reason, a sentence containing only one entity in the source text cannot be used as the target text. For a sentence with 3 or more entities, entity pairs can be formed through pairwise pairing permutation and combination, and can also be selected as the target text. In this way, some meaningless sentences can be filtered out, which helps to reduce the computational amount in subsequent steps and improve efficiency.

[0078] Step 320, label the first entity, the second entity, and the keyword in the target text to obtain the labeled target text; among them, the keyword refers to the words in the target text that can reflect the relationship between the first entity and the second entity.

[0079] The keyword is a word selected from the target text, and the number of keywords can be one or more. The keyword refers to the word that helps to determine the relationship between the first entity and the second entity. The selection of keywords is related to the type of relationship between the entities to be determined. For example, taking the determination of character relationships as an example, assume the target text is "Zhang San is Li Si's first love and Wang Wu's friend", and "first love" and "friend" can be selected as keywords.

[0080] Similarly, when extracting entity relationships, keywords are required. In this regard, sentences without keywords in the source text cannot be used as target texts. For sentences with two or more keywords, they can be selected as target texts.

[0081] In some embodiments, the first entity, the second entity, and keywords in the target text can be marked. Identifiers can be added at the beginning and end of the words to be marked (such as the first entity, the second entity, and keywords) to mark the words. For words in the target text that do not need to be marked, identifiers will not be added at the beginning and end of the words. Among them, the identifier is a symbol used to distinguish / specify the words to be marked. For example, a start identifier is added at the beginning of a word (i.e., in front of the first character of the word), and an end identifier is added at the end of the word (i.e., after the last character of the word). The word located between the start identifier and the end identifier is the marked word. Among them, the start identifier and the end identifier can be the same or different. In the embodiments of the present application, the specific forms of the start identifier and the end identifier are not limited. For example, the start identifier is [X], and the end identifier is [ / X].

[0082] It should be noted that each set of start identifier and end identifier is used to mark an entity or a keyword. When the target text contains multiple entities and / or multiple keywords, each entity has its corresponding set of start identifier and end identifier, and each keyword also has its corresponding set of start identifier and end identifier.

[0083] In some embodiments, the identifiers used to mark entities (including the first entity and the second entity) and keywords are not distinguished. That is, the identifiers corresponding to entities and the identifiers corresponding to keywords are the same.

[0084] In some embodiments, identifiers for annotating entities (including a first entity and a second entity) and keywords are distinguished. That is, the identifiers corresponding to entities and the identifiers corresponding to keywords are different, so that the relation extraction model can determine which word is an entity and which word is a keyword according to the identifier. Exemplarily, the identifier for annotating an entity (including a first entity and a second entity) is an entity identifier, and the entity identifier can also include an entity start identifier and an entity end identifier. The entity start identifier is added in front of the first character of the entity to be annotated, and the entity end identifier is added after the last character of the entity to be annotated. The words between the entity start identifier and the entity end identifier are the entity to be annotated. Similarly, the identifier for annotating a keyword is a keyword identifier, and the keyword identifier can also include a keyword start identifier and a keyword end identifier. The keyword start identifier is added in front of the first character of the keyword to be annotated, and the keyword end identifier is added after the last character of the keyword to be annotated. The words between the keyword start identifier and the keyword end identifier are the keyword to be annotated. Exemplarily, the entity identifier is [PER] and [ / PER], where [PER] is the entity start identifier and [ / PER] is the entity end identifier; the keyword identifier is [Keyword] and [ / Keyword], where [Keyword] is the keyword start identifier and [ / Keyword] is the keyword end identifier.

[0085] In some embodiments, the identifiers for annotating the first entity and the second entity are not distinguished. That is, the identifiers corresponding to the first entity and the identifiers corresponding to the second entity are the same.

[0086] In some embodiments, the identifiers used to label the first entity and the second entity are distinct. That is, the identifier corresponding to the first entity and the identifier corresponding to the second entity are different, so that the relation extraction model can determine which word is the first entity and which word is the second entity based on this identifier. Exemplarily, the identifier used to label the first entity is the first entity identifier, and the first entity identifier can also include a first entity start identifier and a first entity end identifier. The first entity start identifier is added in front of the first word of the labeled first entity, and the first entity end identifier is added behind the last word of the labeled first entity. The words located between the first entity start identifier and the first entity end identifier are the labeled first entity. Similarly, the identifier used to label the second entity is the second entity identifier, and the second entity identifier can also include a second entity start identifier and a second entity end identifier. The second entity start identifier is added in front of the first word of the labeled second entity, and the second entity end identifier is added behind the last word of the labeled second entity. The words located between the second entity start identifier and the second entity end identifier are the labeled second entity. Exemplarily, the first entity identifier is [S:PER] and [ / S:PER], where [S:PER] is the first entity start identifier and [ / S:PER] is the first entity end identifier; the second entity identifier is [O:PER] and [ / O:PER], where [O:PER] is the second entity start identifier and [ / O:PER] is the second entity end identifier.

[0087] In some embodiments, a keyword is a word that determines the relationship between the first entity and the second entity. Since the number of keywords is limited in a specific scenario, the recognition of keywords can be set by enumeration. For example, in the extraction of human relationships, for the human relationship "wife", since the keywords that can yield the relationship of wife are limited, the keywords representing the human relationship of wife can be directly listed in the relationship extraction method, such as "husband and wife", "marriage", "lover", "wife", "husband", "lady", "husband", "husband", "wife", "marry", "wed", "couple", etc. Similarly, the keywords corresponding to other human relationships except "wife" can also be listed by the enumeration method, and finally a keyword library related to the extraction of human relationships is formed. After obtaining the target text, the target text is segmented to obtain multiple words, and then the words belonging to the keyword library among the multiple words are determined as keywords.

[0088] In an exemplary embodiment, such as Figure 4As shown, taking the target text "Zhang San is Li Si's first love and Wang Wu's friend" as an example, assuming the first entity is "Zhang San" and the second entity is "Li Si", and the keywords include "first love" and "friend". The first entity identifiers [S:PER] and [ / S:PER] are placed before and after "Zhang San" to mark the subject entity "Zhang San", which is the first entity; the second entity identifiers [O:PER] and [ / O:PER] are placed before and after "Li Si" to mark the object entity "Li Si", which is the second entity; the keyword identifiers [Keyword] and [ / Keyword] are placed before and after "first love" and "friend" respectively to mark that "first love" and "friend" are keywords. Among them, "Wang Wu" is also a person name entity, but it is not the entity object that needs to be extracted for the relationship, so it does not need to be marked with an identifier. After adding these identifiers to the original target text, it becomes the annotated target text. Among them, [CLS] and [SEP] are placed at the beginning and end of the target text to split the text into individual sentences for convenient encoding.

[0089] Step 330, perform encoding processing on the annotated target text to obtain the encoding representation vector corresponding to the first entity and the entity keyword representation vector, as well as the encoding representation vector corresponding to the second entity and the entity keyword representation vector; among them, the encoding representation vector is used to reflect the feature information of the entity, and the entity keyword representation vector is used to reflect the correlation degree between the entity and the keyword.

[0090] In some embodiments, after obtaining the annotated target text, perform encoding processing on the annotated target text through the encoding network of the relationship extraction model to obtain the encoding representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the first entity, as well as the encoding representation vector corresponding to the second entity and the entity keyword representation vector corresponding to the second entity. Among them, the encoding representation vector corresponding to the first entity is used to reflect the feature information of the first entity and is a vector representation after mapping the semantics of the first entity from low dimension to high dimension; similarly, the encoding representation vector corresponding to the second entity is used to reflect the feature information of the second entity and is a vector representation after mapping the semantics of the second entity from low dimension to high dimension. The entity keyword representation vector corresponding to the first entity is used to reflect the correlation degree between the first entity and the keyword and is a vector representation of the correlation degree between the first entity and the keyword; similarly, the entity keyword representation vector corresponding to the second entity is used to reflect the correlation degree between the second entity and the keyword and is a vector representation of the correlation degree between the second entity and the keyword.

[0091] Optionally, the above encoding network is a neural network, such as a BERT (Bidirectional Encoder Representations from Transformers) network. Essentially, BERT learns a good feature representation for words by running self-supervised learning methods on a large amount of corpus. Self-supervised learning refers to supervised learning running on unlabeled data. The network architecture of BERT uses a multi-layer Transformer structure. Its biggest feature is to abandon the traditional Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN). By means of the attention mechanism, the distance between two words at any position is converted into 1, effectively solving the thorny long-term dependence problem in natural language processing.

[0092] In some embodiments, the encoding network can be a Transformer network. The Transformer network is a deep self-attention transformation network, and is also commonly used to refer to all similar deep self-attention transformation network structures. The Transformer network breaks through the limitation that the recurrent neural network model cannot perform parallel computing. Compared with the convolutional neural network, the number of operations required to calculate the correlation between two positions does not increase with the distance. Self-attention can generate a more interpretable model.

[0093] In some embodiments, the labeled target text is converted into a sequence of embedding representations, and then the sequence of embedding representations is input into the encoding network. After the encoding process of the encoding network, a sequence of encoded representation vectors is output. Then, from the sequence of encoded representation vectors, the encoded representation vector corresponding to the first entity and the encoded representation vector corresponding to the second entity can be extracted. Of course, the encoded representation vector corresponding to the keyword can also be extracted.

[0094] In some embodiments, the intermediate layer feature vectors generated by the encoding network during the encoding process are processed by means of the attention mechanism to obtain the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity. For the generation process of the entity keyword representation vector, please refer to the description in the following embodiments.

[0095] Step 340, determine the relationship between the first entity and the second entity according to the encoded representation vector and the entity keyword representation vector corresponding to the first entity, and the encoded representation vector and the entity keyword representation vector corresponding to the second entity.

[0096] In some embodiments, the encoded representation vectors corresponding to the first entity, the encoded representation vectors corresponding to the second entity, the entity keyword representation vectors corresponding to the first entity, and the entity keyword representation vectors corresponding to the second entity are concatenated to obtain a concatenated vector, and then the concatenated vector is input into the classification network of the relation extraction model, and the relation between the first entity and the second entity is output by the classification network.

[0097] The classification network is a neural network for outputting classification results. For example, it may include a fully connected layer and an output layer (such as the output layer may be a softmax layer), and can output the results of multi-classification.

[0098] In some embodiments, the classification network outputs the confidence levels corresponding to multiple candidate relations, and finally, based on the confidence levels corresponding to the multiple candidate relations, the relation between the first entity and the second entity is determined. For example, the candidate relation with the highest confidence level is selected as the relation between the first entity and the second entity. Of course, in some embodiments, after determining the candidate relation with the highest confidence level, it is determined whether the maximum value of the confidence level is greater than or equal to the threshold value. If the maximum value of the confidence level is greater than or equal to the threshold value, the candidate relation with the highest confidence level is used as the relation between the first entity and the second entity. If the maximum value of the confidence level is less than the threshold value, it is considered that the relation between the first entity and the second entity cannot be determined.

[0099] In summary, the technical solution provided by the embodiments of the present application, by annotating the entities and keywords in the target text, after obtaining the annotated target text, performing encoding processing on the annotated target text to obtain the encoded representation vectors corresponding to the entities and the entity keyword representation vectors, thereby obtaining the feature information of the entities in the target text and the correlation degree between the entities and the keywords, and then determining the relation between the entities based on the above information, thus providing a relation extraction method that does not rely on relation templates, overcomes the problem of insufficient generalization of relation templates, the relation extraction method provided by the present application has strong generalization and can provide a more perfect relation extraction result. Moreover, by extracting the entity keyword representation vectors and using the entity keyword representation vectors to represent the correlation degree between the entities and the keywords, the influence of interfering words in the text on relation extraction is effectively reduced, thus effectively solving the problem that Bootstrapping has poor performance in processing texts with complex contexts and improving the accuracy of relation extraction results.

[0100] At the same time, the relation extraction method provided by the present application only focuses on the entities and keywords in the target text, without considering the positions of the entities and keywords in the target text, so it is applicable to most texts, effectively solving the deficiency of Bootstrapping in terms of generality and improving the generalization of the relation extraction method.

[0101] Please refer toFigure 5 , which shows a flowchart of the relationship extraction method provided by another embodiment of the present application. The execution subject of each step of this method can be Figure 1 the model using device 20 in the implementation environment shown in the figure. This method may include the following steps (510-590):

[0102] Step 510, obtain a target text containing a first entity and a second entity.

[0103] For the introduction and description of the first entity, the second entity, and the target text, please refer to the above embodiments, which will not be elaborated here.

[0104] In some embodiments, as Figure 5 shown, step 510 may include the following sub-steps:

[0105] Step 512, obtain a candidate entity set, where the candidate entity set includes multiple entities;

[0106] Step 514, combine the entities in the candidate entity set in pairs to obtain multiple entity pairs;

[0107] Step 516, for a target entity pair containing the first entity and the second entity among the multiple entity pairs, select a target text containing the first entity and the second entity from the material text.

[0108] In one example, still taking the extraction of person relationships as an example, the candidate entity set includes multiple entities such as "Zhang San", "Li Si", and "Wang Wu". Combine the entities in the candidate entity set in pairs to obtain entity pairs composed of "Zhang San" and "Li Si", entity pairs composed of "Zhang San" and "Wang Wu", entity pairs composed of "Li Si" and "Wang Wu", etc. Then, taking the target entity pair as "Zhang San" and "Li Si" as an example, select a sentence containing "Zhang San" and "Li Si" from the material text as the target text. It can be understood that the number of target texts can be one or multiple.

[0109] Step 520, label the first entity, the second entity, and the keyword in the target text to obtain a labeled target text.

[0110] This step is the same as step 320 in the above embodiment. For specific details, please refer to the introduction in the above embodiment, which will not be elaborated here.

[0111] Step 530, perform encoding processing on the labeled target text through the encoding network of the relationship extraction model to obtain an encoding representation vector corresponding to the first entity and an encoding representation vector corresponding to the second entity.

[0112] In some embodiments, the annotated target text can be regarded as a character sequence, which includes a plurality of characters arranged in order. Each character can be an original word or phrase in the target text or an added identifier. For example, if the target text is "Zhang San is Li Si's first love and Wang Wu's friend", after adding annotations, the resulting annotated target text is "[CLS][S:PER]Zhang San[ / S:PER] is [O:PER]Li Si[ / O:PER]'s [Keyword]first love[ / Keyword], Wang Wu's [Keyword]friend[ / Keyword][SEP]". This annotated target text can be regarded as a character sequence, successively including the following characters: [CLS], [S:PER], Zhang San, [ / S:PER], is, [O:PER], Li Si, [ / O:PER],'s, [Keyword], first love, [ / Keyword], Wang Wu, 's, [Keyword], friend, [ / Keyword], [SEP]. Convert each character into a corresponding token embedding, and obtain a token embedding sequence corresponding to the annotated target text. Token embedding is also called word embedding.

[0113] In some embodiments, input the token embedding sequence corresponding to the annotated target text into an encoding network. After the encoding process of the encoding network, output an encoded representation vector sequence. Then, from this encoded representation vector sequence, the encoded representation vector corresponding to the first entity and the encoded representation vector corresponding to the second entity can be extracted.

[0114] In some embodiments, in addition to having a corresponding token embedding, each character also has a corresponding position embedding. The position embedding is determined according to the position of the character in the annotated target text and is used to represent the position information of the character in the target text.

[0115] In some embodiments, in addition to having a corresponding token embedding, each character also has a corresponding segment embedding. The segment embedding is determined according to the sentence position of the target text in the source text and is used to represent the position information of the target text in the source text.

[0116] In some embodiments, such as Figure 6As shown, for each character in the target text with annotations, the token corresponding to the character is embedded and concatenated with the position embedding and / or segment embedding to obtain the embedding representation corresponding to the character. The embedding representations corresponding to each character are integrated to obtain the embedding representation sequence corresponding to the target text with annotations. The embedding representation sequence corresponding to the target text with annotations is input into the encoding network. After being encoded by the encoding network, a sequence of encoded representation vectors is output. Subsequently, from this sequence of encoded representation vectors, the encoded representation vector corresponding to the first entity (denoted as h(S)) and the encoded representation vector corresponding to the second entity (denoted as h(O)) can be extracted.

[0117] Step 540: Obtain the intermediate layer feature vectors of the encoding network.

[0118] In some embodiments, the encoding network includes multiple encoding layers. The output of the previous encoding layer is the input of the next encoding layer. Each encoding layer encodes the input information once, and the result output by the last encoding layer is the sequence of encoded representation vectors corresponding to the target text with annotations.

[0119] For example, the encoding network includes L encoding layers, where L is an integer greater than 1. The input of the first encoding layer is the embedding representation sequence (or token embedding sequence) corresponding to the target text with annotations. This embedding representation sequence (or token embedding sequence) is encoded by the first encoding layer. The output of the first encoding layer can be referred to as the intermediate layer feature vector corresponding to the first encoding layer. The intermediate layer feature vector corresponding to the first encoding layer is input into the second encoding layer. After being encoded by the second encoding layer, the output of the second encoding layer can be referred to as the intermediate layer feature vector corresponding to the second encoding layer. And so on. The output of the (L - 1)th encoding layer can be referred to as the intermediate layer feature vector corresponding to the (L - 1)th encoding layer. It is input into the Lth encoding layer for encoding, and the Lth encoding layer outputs the sequence of encoded representation vectors corresponding to the target text with annotations.

[0120] In some embodiments, when the encoding network includes L encoding layers, the intermediate layer feature vectors of the encoding network obtained in this step can be the intermediate layer feature vector corresponding to any one of the encoding layers from the first encoding layer to the (L - 1)th encoding layer, or the intermediate layer feature vectors corresponding to multiple encoding layers from the first encoding layer to the (L - 1)th encoding layer. If the intermediate layer feature vectors corresponding to multiple encoding layers are obtained, the intermediate layer feature vectors corresponding to these multiple encoding layers can be fused to obtain the fused intermediate layer feature vector, and then the fused intermediate layer feature vector is used for the subsequent step processing.

[0121] Step 550: Process the intermediate layer feature vectors using an attention mechanism to obtain the entity keyword representation vectors corresponding to the first entity and the entity keyword representation vectors corresponding to the second entity.

[0122] The attention mechanism considers different weight parameters for each element of the input, thereby paying more attention to the parts similar to the input elements while suppressing other useless information. The attention mechanism mimics the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external sensations to increase the observation fineness of some areas. For example, when a person's vision processes a picture, it will quickly scan the global image to obtain the target area that needs to be focused on, that is, the attention focus. Then, more attention resources are invested in this area to obtain more detailed information about the target that needs to be focused on and suppress other useless information. Its greatest advantage is that it can consider global connections and local details in one step. Optionally, the attention mechanism selected here can be the self-attention mechanism. The self-attention mechanism is a variant of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlations of data or features. The application of the self-attention mechanism in text mainly solves the long-distance dependence problem by calculating the mutual influence between words.

[0123] In some embodiments, as Figure 5 shown, step 550 may include the following sub-steps:

[0124] Step 552: From the intermediate layer feature vectors, screen out the intermediate feature vectors corresponding to the first entity, the intermediate feature vectors corresponding to the second entity, and the intermediate feature vectors corresponding to the keywords;

[0125] Step 554: Using the first entity as an anchor point, calculate the attention of the entity keyword representation vector corresponding to the first entity with respect to the intermediate feature vectors corresponding to the second entity and the intermediate feature vectors corresponding to the keywords, to obtain the entity keyword representation vector corresponding to the first entity;

[0126] Step 556: Using the second entity as an anchor point, calculate the attention of the entity keyword representation vector corresponding to the second entity with respect to the intermediate feature vectors corresponding to the first entity and the intermediate feature vectors corresponding to the keywords, to obtain the entity keyword representation vector corresponding to the second entity.

[0127] The attention mechanism is different from the encoding process. It is a process of calculating the correlation degree between each input. Therefore, in this application, the entity keyword representation vectors are obtained by using the attention mechanism to obtain the correlation degree between the entity and the keyword.

[0128] First, as Figure 7As shown, the intermediate feature vectors corresponding to the first entity, the intermediate feature vectors corresponding to the second entity, and the intermediate feature vectors corresponding to the keywords obtained from the above step 552 are used as the inputs of the attention mechanism. Among them, the intermediate feature vector corresponding to the first entity is set as a1, the intermediate feature vector corresponding to the second entity is set as a2, and the intermediate feature vector corresponding to the keyword is set as a3.

[0129] Next, multiply the vectors a1, a2, and a3 by three different embedding transformation matrices W q , W k , W v respectively to obtain different vectors q, k, and v. Taking a1 as an example, three vectors q1, k1, and v1 can be obtained. Among them, q represents the query vector, k represents the key vector, and v represents the information extraction vector.

[0130] Then, taking a1 as an example, multiplying the vector q1 by the vector k1 is a process of attention matching. Among them, in order to prevent the value from being too large, a normalization process is required. After multiplying the vector q1 by the vector k1, it is necessary to divide by to obtain α 1.1 . By analogy, α 1.2 , α 1.3 can be obtained. Among them, d is the dimension of q and k. The dimension is usually understood as: "a point is 0-dimensional, a line is 1-dimensional, a plane is 2-dimensional, and a solid is 3-dimensional". Through this process, the inner product vector α 1.i can be obtained.

[0131] Secondly, we perform the softmax function operation on the obtained inner product vector α 1.i . The softmax function value of this element is the ratio of the exponent of this element to the sum of the exponents of all elements. Taking α 1.1 as an example, its softmax function operation is to divide the exponent of α 1.1 by the exponent of α 1.1 , the exponent of α 1.2 , and the sum of the exponents of α 1.3 . Let it be the value after the softmax function operation for α 1.i .

[0132] After that, multiply the obtained by v i . Specifically, multiply by v1, multiply by v2, multiply Multiply by v3 and add the resulting values to obtain b1, where b1 is the final output result. By analogy, b2 and b3 can be obtained. Among them, b1 is the entity keyword representation vector corresponding to the first entity, and b2 is the entity keyword representation vector corresponding to the second entity.

[0133] Exemplarily, as Figure 6 shown, obtain the intermediate layer feature vector output by the (L-1)-th encoding layer of the encoding network, and use the self-attention mechanism to process this intermediate layer feature vector to obtain the entity keyword representation vector corresponding to the first entity (denoted as h kw (S)) and the entity keyword representation vector corresponding to the second entity (denoted as h kw (O)).

[0134] Step 560: Obtain the difference representation vector between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity.

[0135] Among them, the difference representation vector is used to characterize the difference information between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity.

[0136] In some embodiments, as Figure 5 shown, step 560 may include the following sub-steps:

[0137] Step 562: Subtract the entity keyword representation vector corresponding to the second entity from the entity keyword representation vector corresponding to the first entity to obtain the first difference vector;

[0138] Step 564: Subtract the entity keyword representation vector corresponding to the first entity from the entity keyword representation vector corresponding to the second entity to obtain the second difference vector;

[0139] Step 566: Concatenate the first difference vector and the second difference vector to obtain the difference representation vector.

[0140] Exemplarily, the difference representation vector k is calculated through the following formula diff :

[0141]

[0142] Among them, h kw (S) is the entity keyword representation vector corresponding to the first entity, h kw (O) is the entity keyword representation vector corresponding to the second entity, h kw (S) - h kw (O) is the first difference vector, h kw (O) - hkw(S) is the second difference vector, denotes the concatenation of vectors.

[0143] Step 570: Concatenate the encoded representation vectors corresponding to the first entity, the encoded representation vectors corresponding to the second entity, the entity keyword representation vectors corresponding to the first entity, the entity keyword representation vectors corresponding to the second entity, and the difference representation vector to obtain a concatenated vector.

[0144] Exemplarily, the concatenated vector h is obtained through the following formula kw :

[0145]

[0146] where h(S) is the encoded representation vector corresponding to the first entity, h(O) is the encoded representation vector corresponding to the second entity, h kw (S) is the entity keyword representation vector corresponding to the first entity, h kw (O) is the entity keyword representation vector corresponding to the second entity, kdiff is the difference representation vector, denotes the concatenation of the representation vectors.

[0147] As Figure 6 shown, concatenate the encoded representation vector h(S) corresponding to the first entity, the encoded representation vector h(O) corresponding to the second entity, the entity keyword representation vector h kw (S) corresponding to the first entity, the entity keyword representation vector h kw (O) corresponding to the second entity, and the difference representation vector k diff to obtain a concatenated vector.

[0148] It should be noted that the above step 560 is an optional step. In some embodiments, step 560 may not be executed, and directly concatenate the encoded representation vectors corresponding to the first entity, the encoded representation vectors corresponding to the second entity, the entity keyword representation vectors corresponding to the first entity, and the entity keyword representation vectors corresponding to the second entity to obtain a concatenated vector.

[0149] Step 580: Process the concatenated vector through the classification network of the relation extraction model, and output the confidence levels corresponding to multiple candidate relations.

[0150] As Figure 6 shown, input the above concatenated vector into the classification network, and output the confidence levels corresponding to multiple candidate relations through the classification network. Among them, the classification network includes a fully connected layer and an output layer (such as a softmax layer). The fully connected layer processes the concatenated vector, and then calculates the confidence levels corresponding to multiple candidate relations through the softmax layer.

[0151] Taking character relationship extraction as an example, n candidate relationships can be set, such as spouse, father and daughter, brothers, etc., where n is an integer greater than 1. The confidence value corresponding to each candidate relationship is between [0, 1], and the sum of the confidence values ​​corresponding to the n candidate relationships is 1.

[0152] Step 590: Determine the relationship between the first entity and the second entity based on the confidences corresponding to the plurality of candidate relationships.

[0153] In some embodiments, Figure 5 As shown, step 590 may include the following sub-steps:

[0154] Step 592, selecting a target candidate relationship with the highest confidence according to the confidences corresponding to the multiple candidate relationships;

[0155] Step 594: If the target candidate relationship satisfies the condition, the target candidate relationship is determined to be the relationship between the first entity and the second entity.

[0156] The above conditions include but are not limited to at least one of the following:

[0157] (1) The target text contains words in the whitelist corresponding to the target candidate relationship, and / or the target text does not contain words in the blacklist corresponding to the target candidate relationship;

[0158] In an exemplary embodiment, for the relationship name "spouse", a blacklist and a whitelist are set as shown in Table 1 below. Among them, the blacklist includes words such as "divorce", "proposal", etc., and the whitelist includes words such as "couple", "marriage", "lover", "wife", "husband", "wife", "husband", "husband", "wife", "marry", "marry", "couple", etc. If the target candidate relationship obtained is "spouse", in one example, it is detected whether the target text contains words in the blacklist corresponding to "spouse". If it is contained (for example, the target text contains the word "proposal"), it is determined that the target candidate relationship "spouse" does not meet the conditions, and "spouse" cannot be determined as the relationship between the first entity and the second entity. If it is not contained, it is determined that the target candidate relationship "spouse" meets the conditions. In another example, it is detected whether the target text contains words in the whitelist corresponding to "spouse". If it is contained (for example, the target text contains any word in the whitelist, such as the target text contains the word "lover"), it is determined that the target candidate relationship "spouse" meets the conditions. If it is not contained (for example, any word in the whitelist does not exist in the target text), "spouse" cannot be determined as the relationship between the first entity and the second entity. Of course, in some other examples, it is also possible to determine that the target candidate relationship meets the conditions when the target text contains words in the whitelist corresponding to the target candidate relationship and the target text does not contain words in the blacklist corresponding to the target candidate relationship.

[0159] Table 1

[0160]

[0161]

[0162] (2) The confidence level corresponding to the target candidate relationship is greater than or equal to the first threshold;

[0163] In an exemplary embodiment, assume that the target candidate relationship with the highest confidence level is "spouse", and the first threshold is set to 90%. If the confidence level corresponding to this target candidate relationship is 91%, the relationship between the first entity and the second entity is determined to be "spouse"; if the confidence level corresponding to this target candidate relationship is 87%, the relationship between the entities cannot be determined.

[0164] (3) The number of occurrences of the first entity, the second entity, and the target candidate relationship in the material text is greater than or equal to the second threshold.

[0165] In an exemplary embodiment, assume that the second threshold is set to 20 times. For the material text, multiple target texts containing the first entity and the second entity can be obtained. By performing relationship extraction on the multiple target texts in sequence, multiple extraction results (i.e., the relationship between the first entity and the second entity) can be obtained. Each extraction result may occur multiple times. For example, the number of occurrences of the relationship "spouse" between the first entity and the second entity may be 5 times, 10 times, 15 times, etc. As shown in Table 2 below, the relationship between the characters that can be extracted as "spouse" in two target texts. In this way, when the number of occurrences of the first entity, the second entity, and the target candidate relationship in the material text is greater than or equal to the second threshold, the target candidate relationship is determined as the relationship between the first entity and the second entity, which can retain the high-frequency entity relationships in the material text and discard the low-frequency and rare entity relationships, fully ensuring the accuracy of the final result.

[0166] Table 2

[0167] The first entity Relationship between people The second entity Target text Tom Spouse Lily Tom is Lily's husband. Tom Spouse Lily Tom and Lily take their son to the park.

[0168] In summary, the technical solution provided by the embodiments of the present application enables the first entity and the second entity to ignore the interfering words in the target text, such as verbs and adverbs, through the self-attention mechanism, so that the first entity and the second entity only pay attention to the entities and keywords through the self-attention mechanism, effectively reducing the impact of the interfering words in the text on relationship extraction, thereby effectively solving the problem that Bootstrapping has poor performance in processing texts with complex contexts and improving the accuracy of relationship extraction results.

[0169] In addition, by annotating the first entity, the second entity, and the keywords in the target text, and adding start identifiers and end identifiers before and after the first entity, the second entity, and the keywords, the encoding network can distinguish the entities and keywords in the target text, thereby encoding the target text more accurately and improving the accuracy of the relationship extraction result.

[0170] In addition, by adding a difference representation vector between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity during the encoding process, the influence of the difference between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity on the relationship extraction model is effectively reduced, enabling the classification network to obtain more accurate target candidate relationships when processing the concatenated vector.

[0171] In addition, after obtaining the target candidate relationship with the highest confidence, post-processing is performed. When the target candidate relationship meets the conditions, the target candidate relationship is determined as the relationship between the first entity and the second entity, which can discard the target candidate relationships that do not meet the conditions and retain the target candidate relationships that meet the conditions, ensuring the accuracy of the relationship extraction result.

[0172] Next, Figure 8 , taking the extraction of relationships between characters in film and television dramas as an example, the technical solution of this application will be introduced and explained. Obtain the film and television character vocabulary and film and television text of a certain film and television drama. The film and television character vocabulary contains each character in the film and television drama, and the film and television text is used as the source text, which can include the script and / or lines of the film and television drama. Based on the film and television character vocabulary, multiple pairs of character entities can be constructed, and each pair of character entities contains a first character entity and a second character entity. For each pair of character entities, select the target text containing the pair of character entities from the film and television text. Then, use the relationship extraction method introduced above to extract the relationship between the first character entity and the second character entity from the target text through the relationship extraction model. In some embodiments, post-processing can be performed on the extraction result of the relationship extraction model, such as the black and white list filtering and threshold frequency filtering introduced above, and finally multiple groups of character relationships in the film and television drama are obtained. The multiple groups of character relationships in the film and television drama can be presented to the user in the form of character relationship cards on the client side.

[0173] After a large number of data experiments, using the technical solution of this application in related film and television drama character relationship extraction projects, a total of more than 15,000 character relationships have been extracted, which is about 2.3 times the number of relationships extracted using related technologies. After manual evaluation, the extraction accuracy rate using the technical solution of this application can reach 95.85%, and the character relationship coverage rate can reach 37.14%, proving that this relationship extraction method has the characteristics of a large total number of relationship extractions, high accuracy, and high character relationship coverage rate.

[0174] Next, the training process of the relation extraction model will be introduced through examples. The content involved in the use process of the relation extraction model corresponds to the content involved in the training process, and the two are interconnected. Wherever it is not described in detail on one side, the description on the other side can be referred to.

[0175] Please refer to Figure 9 , which shows the flowchart of the training method of the relation extraction model provided by an embodiment of the present application. The execution subject of each step of this method can be Figure 1 the model training device 10 in the solution implementation environment shown. This method may include the following steps (910-950):

[0176] Step 910, obtain the training samples of the relation extraction model. The training samples include: sample texts containing a first entity and a second entity, and the true relationship between the first entity and the second entity.

[0177] The number of training samples is usually multiple. The true relationship between the first entity and the second entity refers to the determined and accurate relationship between entities, and this true relationship can be the result of manual verification or annotation.

[0178] In addition, the training samples can include positive samples and negative samples. Among them, a positive sample refers to a training sample in which there is a determined true relationship between the first entity and the second entity. For example, if a sentence contains two entities and there is a determined relationship such as "spouse" or "child" between these two entities, then this sentence can be determined as a positive sample. A negative sample refers to a training sample in which there is no determined true relationship between the first entity and the second entity. For example, if a sentence contains two entities and there is no determined relationship such as "spouse" or "child" between these two entities, then this sentence can be determined as a negative sample.

[0179] In some embodiments, the ratio of positive samples to negative samples should be kept balanced to enable the model to have a better training effect.

[0180] Step 920, annotate the first entity, the second entity, and the keyword in the sample text to obtain the annotated sample text. Among them, the keyword refers to the words in the target text that can reflect the relationship between the first entity and the second entity.

[0181] Step 930, perform encoding processing on the annotated sample text through the encoding network of the relation extraction model to obtain the encoding representation vector corresponding to the first entity and the entity keyword representation vector, and the encoding representation vector corresponding to the second entity and the entity keyword representation vector.

[0182] Optionally, step 930 may include the following sub-steps:

[0183] 1. The encoded sample text with annotations is processed through an encoding network to obtain the encoded representation vectors corresponding to the first entity and the second entity;

[0184] 2. Obtain the intermediate layer feature vectors of the encoding network;

[0185] 3. The intermediate layer feature vectors are processed using an attention mechanism to obtain the entity keyword representation vectors corresponding to the first entity and the second entity.

[0186] Step 940, based on the encoded representation vector and the entity keyword representation vector corresponding to the first entity, and the encoded representation vector and the entity keyword representation vector corresponding to the second entity, the classification network of the relation extraction model determines the predicted relationship between the first entity and the second entity.

[0187] Optionally, obtain the difference representation vector between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity. The difference representation vector is used to combine the encoded representation vector and the entity keyword representation vector corresponding to the first entity, and the encoded representation vector and the entity keyword representation vector corresponding to the second entity to determine the predicted relationship between the first entity and the second entity.

[0188] Step 950, based on the true relationship and the predicted relationship, determine the training loss of the relation extraction model, and adjust the network parameters of the relation extraction model based on the training loss.

[0189] The training loss of the relation extraction model is used to measure the prediction accuracy of the model. Optionally, the model parameters are adjusted using the gradient descent method based on the training loss, and finally the relation extraction model that has completed training is obtained.

[0190] For the details not described in detail in the embodiments of the present application, reference can be made to the introduction in the embodiments of the above-mentioned relation extraction method, which will not be elaborated here.

[0191] In summary, the technical solution provided by the embodiments of the present application trains a relation extraction model to implement relation extraction by the model. Compared with the relation extraction method that relies on relation templates, it overcomes the problem of insufficient generalization of relation templates. The relation extraction method provided by the present application has strong generalization and can provide more perfect relation extraction results. Moreover, by extracting the entity keyword representation vectors and using these vectors to characterize the association degree between entities and keywords, the influence of interfering words in the text on relation extraction is effectively reduced, thus effectively solving the problem that Bootstrapping has poor performance in processing texts with complex contexts and improving the accuracy of relation extraction results.

[0192] Please refer to Figure 10, which shows the architecture diagram of the relationship extraction model training.

[0193] First, pre-train the relationship extraction model. Screen out the corpus containing person relationships from the general corpus, extract the entities in the person relationship corpus through named entity recognition, find the positions of the entities in the knowledge graph through entity linking (EL for short), pair the relationships of entity pairs two by two. If there is a relationship between the entity pairs, label the entity pair relationship, mark the entity pair as a positive example, and perform data noise reduction. If there is no relationship between the entity pairs, label the entity pair as a negative example, and balance the number of positive and negative entity pairs. Finally, use the above positive and negative samples to pre-train the relationship extraction model.

[0194] After the pre-training is completed, use the film and television text to construct training samples, and retrain the pre-trained relationship extraction model to finally obtain a relationship extraction model with better performance for film and television person relationship extraction.

[0195] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0196] Please refer to Figure 11 , which shows the block diagram of the relationship extraction device provided by an embodiment of the present application. This device has the function of implementing the above relationship extraction method. The function can be implemented by hardware or by hardware executing corresponding software. This device can be the model usage device introduced above or can be set in the model usage device. The device 1100 may include: a text acquisition module 1110, a labeling module 1120, an encoding processing module 1130, and a relationship determination module 1140.

[0197] The text acquisition module 1110 is used to acquire the target text containing the first entity and the second entity.

[0198] The labeling module 1120 is used to label the first entity, the second entity, and the keyword in the target text to obtain the labeled target text; wherein, the keyword refers to the words in the target text that can reflect the relationship between the first entity and the second entity.

[0199] The encoding processing module 1130 is used to perform encoding processing on the labeled target text to obtain the encoding representation vector corresponding to the first entity and the entity keyword representation vector, as well as the encoding representation vector corresponding to the second entity and the entity keyword representation vector; wherein, the encoding representation vector is used to reflect the feature information of the entity, and the entity keyword representation vector is used to reflect the correlation degree between the entity and the keyword.

[0200] A relationship determination module 1140, configured to determine the relationship between a first entity and a second entity according to the encoded representation vector corresponding to the first entity, the entity keyword representation vector, the encoded representation vector corresponding to the second entity, and the entity keyword representation vector.

[0201] In an exemplary embodiment, an encoding processing module 1130 is configured to:

[0202] Perform encoding processing on the annotated target text through an encoding network to obtain an encoded representation vector corresponding to the first entity and an encoded representation vector corresponding to the second entity;

[0203] Obtain the intermediate layer feature vector of the encoding network;

[0204] Process the intermediate layer feature vector by using an attention mechanism to obtain an entity keyword representation vector corresponding to the first entity and an entity keyword representation vector corresponding to the second entity.

[0205] In an exemplary embodiment, an encoding processing module 1130 is configured to:

[0206] Screen out the intermediate feature vector corresponding to the first entity, the intermediate feature vector corresponding to the second entity, and the intermediate feature vector corresponding to the keyword from the intermediate layer feature vector;

[0207] Taking the first entity as an anchor point, calculate the attention of the intermediate feature vector corresponding to the first entity relative to the intermediate feature vector corresponding to the second entity and the intermediate feature vector corresponding to the keyword, to obtain an entity keyword representation vector corresponding to the first entity;

[0208] Taking the second entity as an anchor point, calculate the attention of the intermediate feature vector corresponding to the second entity relative to the intermediate feature vector corresponding to the first entity and the intermediate feature vector corresponding to the keyword, to obtain an entity keyword representation vector corresponding to the second entity.

[0209] In an exemplary embodiment, the encoding processing module 1130 is further configured to obtain a difference representation vector between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity;

[0210] Wherein, the difference representation vector is used to determine the relationship between the first entity and the second entity by combining the encoded representation vector corresponding to the first entity, the entity keyword representation vector, the encoded representation vector corresponding to the second entity, and the entity keyword representation vector.

[0211] In an exemplary embodiment, an encoding processing module 1130 is configured to:

[0212] Subtract the entity keyword representation vector corresponding to the second entity from the entity keyword representation vector corresponding to the first entity to obtain a first difference vector;

[0213] Subtract the entity keyword representation vector corresponding to the second entity from the entity keyword representation vector corresponding to the first entity to obtain a second difference vector;

[0214] Concatenate the first difference vector and the second difference vector to obtain a difference representation vector.

[0215] In an exemplary embodiment, the relationship determination module 1140 is configured to:

[0216] Concatenate the encoded representation vector corresponding to the first entity, the encoded representation vector corresponding to the second entity, the entity keyword representation vector corresponding to the first entity, and the entity keyword representation vector corresponding to the second entity to obtain a concatenated vector;

[0217] Process the concatenated vector through a classification network and output the confidence levels corresponding to multiple candidate relationships;

[0218] Determine the relationship between the first entity and the second entity based on the confidence levels corresponding to the multiple candidate relationships.

[0219] In an exemplary embodiment, the relationship determination module 1140 is configured to:

[0220] Select a target candidate relationship with the highest confidence level according to the confidence levels corresponding to the multiple candidate relationships;

[0221] If the target candidate relationship meets the conditions, determine the target candidate relationship as the relationship between the first entity and the second entity;

[0222] Wherein, the conditions include at least one of the following:

[0223] The target text contains words in the whitelist corresponding to the target candidate relationship, and / or, the target text does not contain words in the blacklist corresponding to the target candidate relationship;

[0224] The confidence level corresponding to the target candidate relationship is greater than or equal to a first threshold;

[0225] The number of occurrences of the first entity, the second entity, and the target candidate relationship in the material text is greater than or equal to a second threshold.

[0226] In an exemplary embodiment, the text acquisition module 1110 is configured to:

[0227] Obtain a candidate entity set, where the candidate entity set includes multiple entities;

[0228] Combine the entities in the candidate entity set in pairs to obtain multiple entity pairs;

[0229] For a target entity pair containing a first entity and a second entity among multiple entity pairs, select a target text containing the first entity and the second entity from the material text.

[0230] In summary, the technical solution provided by the embodiments of the present application, by annotating the entities and keywords in the target text, after obtaining the annotated target text, performing encoding processing on the annotated target text to obtain the encoded representation vectors corresponding to the entities and the entity keyword representation vectors, thereby obtaining the feature information of the entities in the target text and the correlation degree between the entities and the keywords, and then determining the relationship between the entities based on the above information, thus providing a relationship extraction method that does not rely on a relationship template, overcoming the problem of insufficient generalization of the relationship template. The relationship extraction method provided by the present application has strong generalization and can provide a more perfect relationship extraction result. Moreover, by extracting the entity keyword representation vectors and using the entity keyword representation vectors to characterize the correlation degree between the entities and the keywords, the influence of interfering words in the text on relationship extraction is effectively reduced, thereby effectively solving the problem that Bootstrapping has poor performance in processing texts with complex contexts and improving the accuracy of the relationship extraction result.

[0231] At the same time, the relationship extraction method provided by the present application only focuses on the entities and keywords in the target text and does not need to consider the positions of the entities and keywords in the target text, so it is applicable to most texts, effectively solving the deficiency of Bootstrapping in terms of generality and improving the generalization of the relationship extraction method.

[0232] Please refer to Figure 12 , which shows a block diagram of a training device for a relationship extraction model provided by an embodiment of the present application. This device has the function of implementing the training method of the above-mentioned relationship extraction model, and this function can be implemented by hardware or by hardware executing corresponding software. This device can be the model training device introduced above or can be set in the model training device. The device 1200 may include: a sample acquisition module 1210, a sample annotation module 1220, an encoding processing module 1230, a relationship determination module 1240, and a parameter adjustment module 1250.

[0233] The sample acquisition module 1210 is used to acquire training samples for the relationship extraction model. The training samples include: sample texts containing a first entity and a second entity, and the true relationship between the first entity and the second entity.

[0234] The sample annotation module 1220 is used to annotate the first entity, the second entity, and the keywords in the sample text to obtain an annotated sample text; wherein, the keywords refer to the words and phrases in the sample text that can reflect the relationship between the first entity and the second entity.

[0235] The encoding processing module 1230 is configured to perform encoding processing on the annotated sample text through the encoding network of the relation extraction model to obtain the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, as well as the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors; wherein, the encoded representation vectors are used to reflect the feature information of the entity, and the entity keyword representation vectors are used to reflect the correlation degree between the entity and the keyword.

[0236] The relation determination module 1240 is configured to determine the predicted relation between the first entity and the second entity through the classification network of the relation extraction model according to the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, as well as the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors.

[0237] The parameter adjustment module 1250 is configured to determine the training loss of the relation extraction model according to the true relation and the predicted relation, and adjust the network parameters of the relation extraction model based on the training loss.

[0238] In an exemplary embodiment, the encoding processing module 1230 is configured to:

[0239] Perform encoding processing on the annotated sample text through the encoding network to obtain the encoded representation vectors corresponding to the first entity and the encoded representation vectors corresponding to the second entity;

[0240] Obtain the intermediate layer feature vectors of the encoding network;

[0241] Process the intermediate layer feature vectors by using the attention mechanism to obtain the entity keyword representation vectors corresponding to the first entity and the entity keyword representation vectors corresponding to the second entity.

[0242] In an exemplary embodiment, the relation determination module 1240 is configured to:

[0243] Obtain the difference representation vector between the entity keyword representation vectors corresponding to the first entity and the entity keyword representation vectors corresponding to the second entity;

[0244] Wherein, the difference representation vector is used to combine the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors and the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors to determine the predicted relation between the first entity and the second entity.

[0245] In summary, for the technical solution provided in the embodiments of the present application, by training a relation extraction model to implement relation extraction by the model, compared with the relation extraction method that relies on relation templates, the problem of insufficient generalization of relation templates is overcome. The relation extraction method provided in the present application has strong generalization ability and can provide more perfect relation extraction results. Moreover, by extracting the entity keyword representation vectors and using these vectors to characterize the association degree between the entity and the keyword, the influence of interfering words in the text on relation extraction is effectively reduced, thus effectively solving the problem that Bootstrapping has poor effect in processing texts with complex contexts and improving the accuracy of relation extraction results.

[0246] It should be noted that, when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.

[0247] Please refer to Figure 13 , which shows the structural schematic diagram of a computer device provided in an embodiment of the present application. The computer device can be any electronic device with data calculation, processing, and storage functions, such as a mobile phone, a tablet computer, a PC (Personal Computer), or a server, etc. The computer device can be implemented as a model usage device for implementing the relation extraction method provided in the above embodiments; or, the computer device can be implemented as a model training device for implementing the training method of the relation extraction model provided in the above embodiments. Specifically:

[0248] The computer device 1100 includes a processing unit (such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array), etc.) 1301, a system memory 1304 including a RAM (Random-Access Memory) 1302 and a ROM (Read-Only Memory) 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The computer device 1300 also includes a basic input / output system (Input Output System, I / O system) 1306 for facilitating the transfer of information between various components within the server, and a mass storage device 1307 for storing an operating system 1313, application programs 1314, and other program modules 1315.

[0249] The basic input / output system 1306 includes a display 1308 for displaying information and input devices such as a mouse, a keyboard, etc. 1309 for user input of information. Among them, both the display 1308 and the input devices 1309 are connected to the central processing unit 1301 through an input / output controller 1310 connected to the system bus 1305. The basic input / output system 1306 may also include an input / output controller 1310 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1310 also provides outputs to a display screen, a printer, or other types of output devices.

[0250] The mass storage device 1307 is connected to the central processing unit 1301 through a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1307 and its associated computer-readable medium provide non-volatile storage for the computer device 1300. That is to say, the mass storage device 1307 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.

[0251] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media is not limited to the above several types. The above-mentioned system memory 1304 and mass storage device 1307 can be collectively referred to as memory.

[0252] According to an embodiment of the present application, the computer device 1300 can also be run by a remote computer on the network through a network such as the Internet. That is, the computer device 1300 can be connected to the network 1312 through the network interface unit 1311 connected to the system bus 1305, or in other words, the network interface unit 1311 can also be used to connect to other types of networks or remote computer systems (not shown).

[0253] The memory further includes at least one instruction, at least one program, a code set, or an instruction set, which is stored in the memory and is configured to be executed by one or more processors to implement the above-mentioned relation extraction method or the training method of the relation extraction model.

[0254] In an exemplary embodiment, a computer-readable storage medium is also provided. The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and when the at least one instruction, at least one program, a code set, or an instruction set is executed by the processor of the computer device, it implements the relation extraction method or the training method of the relation extraction model provided in the above embodiment.

[0255] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0256] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned entity extraction method or the training method of the relationship extraction model.

[0257] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. In addition, the step numbers described herein only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.

[0258] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A relation extraction method, characterized in that, The method includes: Obtaining a target text including a first entity and a second entity; Annotating the first entity, the second entity, and keywords in the target text to obtain an annotated target text; wherein, the keywords refer to words in the target text that can reflect the relationship between the first entity and the second entity; Performing encoding processing on the annotated target text through an encoding network to obtain an encoding representation vector corresponding to the first entity and an encoding representation vector corresponding to the second entity; Obtaining an intermediate layer feature vector of the encoding network; From the intermediate layer feature vectors, screening out an intermediate feature vector corresponding to the first entity, an intermediate feature vector corresponding to the second entity, and an intermediate feature vector corresponding to the keywords; Taking the first entity as an anchor point, calculating the attention of the intermediate feature vector corresponding to the first entity relative to the intermediate feature vector corresponding to the second entity and the intermediate feature vector corresponding to the keywords to obtain an entity-keyword representation vector corresponding to the first entity; Taking the second entity as an anchor point, calculating the attention of the intermediate feature vector corresponding to the second entity relative to the intermediate feature vector corresponding to the first entity and the intermediate feature vector corresponding to the keywords to obtain an entity-keyword representation vector corresponding to the second entity; Determining the relationship between the first entity and the second entity according to the encoding representation vector and the entity-keyword representation vector corresponding to the first entity, and the encoding representation vector and the entity-keyword representation vector corresponding to the second entity.

2. The method according to claim 1, characterized in that The method further includes: Obtaining a difference representation vector between the entity-keyword representation vector corresponding to the first entity and the entity-keyword representation vector corresponding to the second entity; Wherein, the difference representation vector is used to determine the relationship between the first entity and the second entity by combining the encoding representation vector and the entity-keyword representation vector corresponding to the first entity and the encoding representation vector and the entity-keyword representation vector corresponding to the second entity.

3. The method according to claim 2, wherein The obtaining the difference information between the entity-keyword representation vector corresponding to the first entity and the entity-keyword representation vector corresponding to the second entity includes: Subtracting the entity-keyword representation vector corresponding to the second entity from the entity-keyword representation vector corresponding to the first entity to obtain a first difference vector; Subtracting the entity-keyword representation vector corresponding to the first entity from the entity-keyword representation vector corresponding to the second entity to obtain a second difference vector; Concatenating the first difference vector and the second difference vector to obtain the difference representation vector.

4. The method according to claim 1, wherein The determining the relationship between the first entity and the second entity according to the encoding representation vector and the entity-keyword representation vector corresponding to the first entity, and the encoding representation vector and the entity-keyword representation vector corresponding to the second entity includes: Concatenating the encoding representation vector corresponding to the first entity, the encoding representation vector corresponding to the second entity, the entity-keyword representation vector corresponding to the first entity, and the entity-keyword representation vector corresponding to the second entity to obtain a concatenated vector; Process the spliced vector through a classification network to output the confidence levels corresponding to multiple candidate relationships; Based on the confidence levels corresponding to the multiple candidate relationships, determine the relationship between the first entity and the second entity.

5. The method according to claim 4, wherein The determining the relationship between the first entity and the second entity based on the confidence levels corresponding to the multiple candidate relationships includes: According to the confidence levels corresponding to the multiple candidate relationships, select the target candidate relationship with the highest confidence level; If the target candidate relationship meets the conditions, determine the target candidate relationship as the relationship between the first entity and the second entity; Wherein, the conditions include at least one of the following: The target text contains words in the whitelist corresponding to the target candidate relationship, and / or, the target text does not contain words in the blacklist corresponding to the target candidate relationship; The confidence level corresponding to the target candidate relationship is greater than or equal to a first threshold; The number of occurrences of the first entity, the second entity, and the target candidate relationship in the material text is greater than or equal to a second threshold.

6. The method according to claim 1, characterized in that The obtaining the target text containing the first entity and the second entity includes: Obtain a candidate entity set, where the candidate entity set includes multiple entities; Combine the entities in the candidate entity set in pairs to obtain multiple entity pairs; For the target entity pair containing the first entity and the second entity among the multiple entity pairs, select the target text containing the first entity and the second entity from the material text.

7. A training method for a relationship extraction model, characterized in that The method includes: Obtain training samples of a relationship extraction model, where the training samples include: a sample text containing a first entity and a second entity, and the true relationship between the first entity and the second entity; Annotate the first entity, the second entity, and keywords in the sample text to obtain an annotated sample text; wherein, the keywords refer to the words in the sample text that can reflect the relationship between the first entity and the second entity; Perform encoding processing on the annotated sample text through an encoding network of the relationship extraction model to obtain an encoding representation vector corresponding to the first entity and an encoding representation vector corresponding to the second entity; Obtain the intermediate layer feature vector of the encoding network; From the intermediate layer feature vector, filter out the intermediate feature vector corresponding to the first entity, the intermediate feature vector corresponding to the second entity, and the intermediate feature vector corresponding to the keyword; Taking the first entity as an anchor point, calculate the attention of the intermediate feature vector corresponding to the first entity relative to the intermediate feature vector corresponding to the second entity and the intermediate feature vector corresponding to the keyword to obtain an entity-keyword representation vector corresponding to the first entity; Taking the second entity as an anchor point, calculate the attention of the intermediate feature vector corresponding to the second entity relative to the intermediate feature vector corresponding to the first entity and the intermediate feature vector corresponding to the keyword to obtain an entity-keyword representation vector corresponding to the second entity; Based on the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors, as well as the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors, the classification network of the relationship extraction model determines the predicted relationship between the first entity and the second entity; Based on the true relationship and the predicted relationship, the training loss of the relationship extraction model is determined, and the network parameters of the relationship extraction model are adjusted based on the training loss.

8. The method according to claim 7, characterized in that, The method further includes: Obtaining a difference representation vector between the entity keyword representation vector corresponding to the first entity and the entity keyword representation vector corresponding to the second entity; Wherein, the difference representation vector is used to determine the predicted relationship between the first entity and the second entity in combination with the encoded representation vectors corresponding to the first entity and the entity keyword representation vectors and the encoded representation vectors corresponding to the second entity and the entity keyword representation vectors.

9. A relation extraction device, characterized in that, The apparatus includes: A text acquisition module, configured to acquire a target text including a first entity and a second entity; A labeling module, configured to label the first entity, the second entity, and keywords in the target text to obtain a labeled target text; wherein, the keywords refer to words in the target text that can reflect the relationship between the first entity and the second entity; An encoding processing module, configured to perform encoding processing on the labeled target text through an encoding network to obtain an encoded representation vector corresponding to the first entity and an encoded representation vector corresponding to the second entity; The encoding processing module is further configured to obtain an intermediate layer feature vector of the encoding network; The encoding processing module is further configured to filter out the intermediate feature vector corresponding to the first entity, the intermediate feature vector corresponding to the second entity, and the intermediate feature vector corresponding to the keyword from the intermediate layer feature vector; The encoding processing module is further configured to calculate the attention of the intermediate feature vector corresponding to the first entity relative to the intermediate feature vector corresponding to the second entity and the intermediate feature vector corresponding to the keyword with the first entity as an anchor point to obtain an entity keyword representation vector corresponding to the first entity; The encoding processing module is further configured to calculate the attention of the intermediate feature vector corresponding to the second entity relative to the intermediate feature vector corresponding to the first entity and the intermediate feature vector corresponding to the keyword with the second entity as an anchor point to obtain an entity keyword representation vector corresponding to the second entity; A relationship determination module, configured to determine the relationship between the first entity and the second entity based on the encoded representation vector corresponding to the first entity and the entity keyword representation vector, as well as the encoded representation vector corresponding to the second entity and the entity keyword representation vector.

10. A training device for a relationship extraction model, characterized in that, The apparatus includes: A sample acquisition module, configured to acquire training samples of a relationship extraction model, where the training samples include: a sample text including a first entity and a second entity, and the true relationship between the first entity and the second entity; A sample annotation module for annotating the first entity, the second entity, and keywords in the sample text to obtain an annotated sample text; wherein, the keywords refer to words in the sample text that can reflect the relationship between the first entity and the second entity; An encoding processing module for encoding the annotated sample text through the encoding network of the relationship extraction model to obtain an encoding representation vector corresponding to the first entity and an encoding representation vector corresponding to the second entity; The encoding processing module is further configured to obtain an intermediate layer feature vector of the encoding network; The encoding processing module is further configured to screen out an intermediate feature vector corresponding to the first entity, an intermediate feature vector corresponding to the second entity, and an intermediate feature vector corresponding to the keyword from the intermediate layer feature vector; The encoding processing module is further configured to calculate the attention of the intermediate feature vector corresponding to the first entity relative to the intermediate feature vector corresponding to the second entity and the intermediate feature vector corresponding to the keyword with the first entity as an anchor point to obtain an entity-keyword representation vector corresponding to the first entity; The encoding processing module is further configured to calculate the attention of the intermediate feature vector corresponding to the second entity relative to the intermediate feature vector corresponding to the first entity and the intermediate feature vector corresponding to the keyword with the second entity as an anchor point to obtain an entity-keyword representation vector corresponding to the second entity; A relationship determination module for determining a predicted relationship between the first entity and the second entity through the classification network of the relationship extraction model according to the encoding representation vector and the entity-keyword representation vector corresponding to the first entity, and the encoding representation vector and the entity-keyword representation vector corresponding to the second entity; A parameter adjustment module for determining a training loss of the relationship extraction model according to the true relationship and the predicted relationship, and adjusting network parameters of the relationship extraction model based on the training loss.

11. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the relationship extraction method according to any one of claims 1 to 6, or to implement the training method of the relationship extraction model according to claim 7 or 8.

12. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the relationship extraction method according to any one of claims 1 to 6, or to implement the training method of the relationship extraction model according to claim 7 or 8.

13. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to execute the relationship extraction method according to any one of claims 1 to 6, or to execute the training method of the relationship extraction model according to claim 7 or 8.

Citation Information

Patent Citations

  • An entity relationship joint extraction method and system based on an attention mechanism

    CN109902145A

  • Entity relationship extraction method of concerned associated words

    CN110196978A