Entity relationship extraction method and device
By obtaining and processing part-of-speech sequences of text and selecting the tag sequence pattern to determine the labels of unlabeled word segmentation, the accuracy and flexibility of entity relationship extraction in the prior art is solved, and efficient and accurate entity relationship extraction is achieved.
Patent Information
- Application Number
- CN202010659025.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-07-09
AI Technical Summary
The existing entity relationship extraction methods have problems such as low accuracy, poor flexibility, and poor transplant performance.
By obtaining the first part-of-speech sequence and the second part-of-speech sequence of multiple texts, the tag sequence pattern is selected to determine the tags of unlabeled word segments in the target text, thereby generating the entity relationship extraction result.
Efficient and accurate physical relationship extraction is achieved, reducing the cost of human maintenance rules and improving the flexibility and accuracy of extraction.
Smart Images

Figure CN111753029B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing and information extraction, and more particularly, to a method and apparatus for entity relationship extraction. Background Art
[0002] Entity relationship extraction is of great significance for portrait construction and graph construction. For example, by mining and extracting entity relationships in corpora such as financial news and forum opinions, enterprise portraits, merchant portraits or industry portraits can be constructed based on entity relationship extraction, thus creating value for applications such as industry analysis and strategic analysis; in social portrait mining and relationship chain construction, by extracting entity relationships between people's names, social relationship chains and person relationship graphs can be constructed, thus enabling applications such as social recommendation and relationship network marketing. However, there are still many problems in existing entity relationship extraction methods, such as low accuracy of extraction results, poor flexibility, and poor portability. Summary of the Invention
[0003] Embodiments of this application provide a method and apparatus for entity relationship extraction, which can at least to some extent achieve efficient and accurate extraction of entity relationships.
[0004] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.
[0005] According to one aspect of the embodiments of this application, a method for entity relationship extraction is provided, including: obtaining a first part-of-speech sequence corresponding to each of a plurality of texts, where the first part-of-speech sequence corresponding to each text includes part-of-speech elements corresponding to each word segment in the word segmentation result of each text; mapping the tags of the word segments with annotations in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text to generate a second part-of-speech sequence corresponding to each text, where the tags include entity tags and entity relationship tags; selecting at least one second part-of-speech sequence from the second part-of-speech sequences corresponding to the plurality of texts as a tag sequence pattern according to the first part-of-speech sequences corresponding to the plurality of texts; and determining the tags of the word segments without annotations in the target text according to the tag sequence pattern, so as to generate an entity relationship extraction result of the target text according to the tags in the target text.
[0006] According to one aspect of the embodiments of the present application, there is provided an entity relationship extraction device, including: an acquisition unit configured to acquire first part-of-speech sequences respectively corresponding to a plurality of texts, where the first part-of-speech sequence corresponding to each text includes part-of-speech elements corresponding to each word segmentation in the word segmentation result of each text; a generation unit configured to map the labels of the segmented words with labels in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text, and generate a second part-of-speech sequence corresponding to each text, where the labels include entity labels and entity relationship labels; a selection unit configured to select at least one second part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the plurality of texts as a label sequence pattern according to the first part-of-speech sequences respectively corresponding to the plurality of texts; a determination unit configured to determine the labels of the segmented words without labels in the target text according to the label sequence pattern, so as to generate an entity relationship extraction result of the target text according to the labels in the target text.
[0007] In some embodiments of the present application, based on the foregoing solution, the selection unit includes: a first selection subunit configured to select part-of-speech elements with a first support degree greater than a first threshold in the plurality of texts from the first part-of-speech sequences respectively corresponding to the plurality of texts, and obtain third part-of-speech sequences respectively corresponding to the plurality of texts; a mining subunit configured to perform sequence pattern mining on the third part-of-speech sequences respectively corresponding to the plurality of texts to generate frequent sequence patterns; a second selection subunit configured to select at least one second part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the plurality of texts as the label sequence pattern according to the frequent sequence patterns.
[0008] In some embodiments of the present application, based on the foregoing solution, the first selection subunit is further configured to: count the number of texts including each part-of-speech element in the plurality of texts according to the part-of-speech elements in the first part-of-speech sequences respectively corresponding to the plurality of texts; calculate a ratio between the number of texts including each part-of-speech element and the total number of the plurality of texts to obtain the first support degree of each part-of-speech element in the plurality of texts.
[0009] In some embodiments of the present application, based on the foregoing solution, the mining subunit is further configured to: select a part-of-speech element as a prefix from the third part-of-speech sequences respectively corresponding to the multiple texts, and determine at least one suffix corresponding to the prefix, where the at least one suffix includes the part-of-speech elements in the third part-of-speech sequence that are located after the prefix, and the order of the included part-of-speech elements is the same as the order in the third part-of-speech sequence; select a part-of-speech element with a second support degree greater than the first threshold in the at least one suffix and add it to the prefix to obtain a new prefix, and continue to determine a new suffix corresponding to the new prefix until no part-of-speech element with a second support degree greater than the threshold can be selected from the determined new suffix; generate the frequent sequence pattern according to the obtained multiple prefixes.
[0010] In some embodiments of the present application, based on the foregoing solution, the mining subunit is further configured to: if there is a target prefix among the multiple prefixes that includes the part-of-speech elements in other prefixes and the order of the included part-of-speech elements is the same as the order in the other prefixes, then use the target prefix as the frequent sequence pattern.
[0011] In some embodiments of the present application, based on the foregoing solution, the mining subunit is further configured to: count the number of suffixes including each part-of-speech element in the at least one suffix according to the part-of-speech elements in the at least one suffix; calculate the ratio between the number of suffixes including each part-of-speech element and the total number of the multiple texts to obtain the second support degree of each part-of-speech element in the at least one suffix.
[0012] In some embodiments of the present application, based on the foregoing solution, the second selection subunit is further configured to: select at least one target part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the multiple texts according to the frequent sequence pattern, where the at least one target part-of-speech sequence includes the part-of-speech elements in the frequent sequence pattern and the order of the included part-of-speech elements is the same as the order in the frequent sequence pattern; calculate the ratio between the number of tags in each target part-of-speech sequence and the sum of the number of tags in the at least one target part-of-speech sequence to obtain the confidence corresponding to each target part-of-speech sequence; use the target part-of-speech sequence with a confidence greater than the second threshold as the tag sequence pattern.
[0013] In some embodiments of the present application, based on the foregoing solution, the second selection subunit is further configured to: obtain the position serial numbers corresponding to the tags in each target part-of-speech sequence in each target part-of-speech sequence; sum the number of tags with different position serial numbers in the at least one target part-of-speech sequence to obtain the sum of the number of tags in the at least one target part-of-speech sequence.
[0014] In some embodiments of the present application, based on the foregoing solution, the determining unit is further configured to: select at least one target tag sequence pattern from the tag sequence patterns according to the first part-of-speech sequence corresponding to the target text, where the at least one target tag sequence pattern includes the part-of-speech elements in the first part-of-speech sequence corresponding to the target text, and the position order of the included part-of-speech elements is the same as the position order in the first part-of-speech sequence corresponding to the target text; determine the tags of the unannotated word segments in the target text according to the at least one target tag sequence pattern.
[0015] In the technical solutions provided by some embodiments of the present application, by performing word segmentation and part-of-speech annotation processing on multiple texts, first part-of-speech sequences respectively corresponding to the multiple texts are obtained, and the tags of the annotated word segments in each text are mapped to the part-of-speech elements of the first part-of-speech sequence to generate a second part-of-speech sequence corresponding to each text. According to the part-of-speech elements in the first part-of-speech sequence, at least one second part-of-speech sequence is selected from the second part-of-speech sequences corresponding to each text as the tag sequence pattern. Finally, the tags of the unannotated word segments in the target text are determined according to the tag sequence pattern, and combined with the tags of the annotated word segments in the target text, the entity relationship extraction result of the target text is generated. Compared with the prior art, the technical solution of the embodiments of the present application generates a tag sequence pattern based on the part-of-speech elements in the text and the tags of the annotated word segments in the text. As the text is updated, the generated tag sequence pattern will also change, so that the extraction of entity relationships does not depend on fixed extraction rules, reducing the cost of manual maintenance of rules. At the same time, the tag sequence pattern changes with the change of the text, considering the actual situation of the text, ensuring the accuracy of the generated tag sequence pattern, thereby improving the accuracy of entity relationship extraction. Moreover, the technical solution of the embodiments of the present application does not require complex network training such as a neural network model, so that the efficiency of entity relationship extraction can be improved, realizing efficient and flexible extraction of entity relationships.
[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0018] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown;
[0019] Figure 2 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0020] Figure 3 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0021] Figure 4 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0022] Figure 5 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0023] Figure 6 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0024] Figure 7 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0025] Figure 8 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0026] Figure 9 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application;
[0027] Figure 10 shows a block diagram of an entity relationship extraction device according to an embodiment of the present application;
[0028] Figure 11 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0029] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0030] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0031] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0032] The flowcharts shown in the drawings are merely illustrative and not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0033] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural.
[0034] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0035] 1) Entity: Something in the real world that is distinguishable and exists independently, such as: person names, place names, game names, etc.
[0036] 2) Relationship extraction: A relationship is defined as the connection between two or more entities. Relationship extraction is to identify the relationship by learning the semantic connection between multiple entities in the text. The input of relationship extraction is a paragraph or a sentence of text, and the output is usually a triple: <entity 1, relationship, entity 2>. For example, for the input text "The composer of song a is singer A", after relationship extraction, it can be output as <song a, composer, singer A>.
[0037] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied is shown.
[0038] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, the network 104, and the server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with display screens, including but not limited to desktop computers, portable computers, smartphones, and tablet computers, etc. It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0039] are merely illustrative, and according to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.
[0040] The technical solutions of the embodiments of the present application will be elaborated in detail below:
[0041] In information extraction technology, entity relationship extraction is a necessary link in portrait construction and atlas construction. Currently, the methods of entity relationship extraction mainly include entity relationship extraction methods based on vocabulary-semantics, entity relationship extraction based on annotated corpus machine learning, and entity relationship extraction methods based on pattern mining and matching.
[0042] Among them, the entity relationship extraction method based on lexicon-semantics first uses the word vector method to extract concept synonyms from the corpus to construct a concept dictionary, then annotates lexical information, syntactic information, and semantic information, and designs a lexical-semantic rule annotation algorithm based on the finite state machine theory for automatic annotation, so as to identify which components in the sentence constitute important elements of the entity relationship. However, the relationship extraction method based on lexicon-semantics depends on the effect of the word vector model, and often introduces some noise words when expanding synonyms, which affects the accuracy of the extraction result.
[0043] The entity relationship extraction method based on machine learning needs to first give candidate event elements, treat the relationship extraction as a classification problem, and determine whether the candidate relationship element is a relationship element of the relationship through classifiers such as Support Vector Machine (SVM). The method of using dependency syntax for entity sentence syntax relationship extraction depends on the accuracy of syntactic analysis and also depends on the accuracy of the word segmentation result. The error accumulation of the two subtasks of word segmentation and syntactic analysis will lead to the error superposition of the parent task relationship extraction, which is often difficult to improve in practical applications.
[0044] The entity relationship extraction method based on pattern matching and mining first needs to annotate the relationship elements and trigger words in the sentence, then construct the relationship between the relationship elements and trigger words in the syntactic tree, and perform pattern mining and construction according to the existing rules of syntactic relationships, so as to extract the relationship element information. The method of relationship elements based on pattern matching and mining has problems such as poor flexibility, low recall rate, and poor portability.
[0045] In view of this, the embodiments of the present application provide an entity relationship extraction method, which can efficiently and accurately extract entity relationships, with good flexibility and operability.
[0046] See Figure 2 , Figure 2 shows a flowchart of an entity relationship extraction method according to an embodiment of the present application. The entity relationship extraction method can be executed by a server, and the server can be Figure 1 the server 105 shown in Figure 1 Of course, the entity relationship extraction method can also be executed by a terminal device, such as can be executed by Figure 2 the terminals 101, 102, 103 shown in
[0047] Step S210: Obtain first part-of-speech sequences corresponding to multiple texts, and each first part-of-speech sequence corresponding to a text includes part-of-speech elements corresponding to each word segment in the word segmentation result of each text;
[0048] Step S220: Map the tags of the segmented words already marked in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text, and generate a second part-of-speech sequence corresponding to each text, where the tags include entity tags and entity relationship tags;
[0049] Step S230: According to the first part-of-speech sequences respectively corresponding to the multiple texts, select at least one second part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the multiple texts as the tag sequence pattern;
[0050] Step S240: Determine the tags of the unmarked segmented words in the target text according to the tag sequence pattern, so as to generate the entity relationship extraction result of the target text based on the tags in the target text.
[0051] The following describes these steps in detail.
[0052] In step S210, the first part-of-speech sequences respectively corresponding to multiple texts are obtained, and the first part-of-speech sequence corresponding to each text includes the part-of-speech elements corresponding to each segmented word in the segmented result of each text.
[0053] After the server obtains multiple texts, it can first perform preprocessing on the multiple texts respectively. The preprocessing includes word segmentation processing and stop word removal processing, and obtains the segmented sequences respectively corresponding to the multiple texts. For stop word removal, for example, punctuation marks can be filtered out; furthermore, for each segmented word in the segmented sequence, part-of-speech tagging processing is performed to obtain the first part-of-speech sequences respectively corresponding to the multiple texts. The first part-of-speech sequence includes the part-of-speech elements corresponding to each segmented word in the segmented result of each text, and each part-of-speech element is arranged in order. For example, for the text "The composer of song a is singer A", the segmented sequence "song a / of / composer / is / singer A" can be obtained through word segmentation processing, and the first part-of-speech sequence " / n / u / n / v / nr" can be obtained through part-of-speech tagging processing. The part-of-speech elements in the first part-of-speech sequence are arranged in order. The "n" in the first position represents the part of speech of "song a", which is a noun, the "u" in the second position represents the part of speech of "of", which is a particle, the "n" in the third position represents the part of speech of "composer", which is a noun, the "v" in the fourth position represents the part of speech of "is", which is a verb, and the "nr" in the fifth position represents the part of speech of "singer A", which is a personal name.
[0054] It should be noted that there are already relatively mature word segmentation processing methods and part-of-speech tagging methods in the related art. Here, the word segmentation processing method in the related art can be directly used to perform word segmentation processing on the target text, and the part-of-speech tagging method in the related art can be used to perform part-of-speech tagging processing on the segmented sequence obtained by the word segmentation processing. The present application does not make any limitation on the specific word segmentation processing method and part-of-speech tagging method adopted.
[0055] Step S220: Map the tags of the segmented words with annotations in each of the texts to the part-of-speech elements of the first part-of-speech sequence corresponding to each of the texts, to generate a second part-of-speech sequence corresponding to each of the texts, where the tags include entity tags and entity relationship tags.
[0056] Specifically, entities may include personal names, place names, organizations, time, numbers, etc.; entity relationships may include social relationships of people, physical orientation relationships, general subordination relationships, whole-part relationships, organizational subordination relationships, relationships of owned items, etc.
[0057] In the embodiments of the present application, among multiple texts, some segmented words are annotated with tags. Among them, the segmented words belonging to entities correspond to entity tags, and the segmented words belonging to entity relationships correspond to entity relationship tags. After performing word segmentation processing and part-of-speech tagging processing on multiple texts respectively to obtain the first part-of-speech sequence, the tags of the segmented words with annotations can be mapped to the part-of-speech elements of the first part-of-speech sequence corresponding to each text, so as to obtain the second part-of-speech sequence.
[0058] Step S230: According to the first part-of-speech sequences respectively corresponding to the multiple texts, select at least one second part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the multiple texts as the tag sequence pattern.
[0059] In the embodiments of the present application, the tag sequence pattern can be obtained by selecting from the second part-of-speech sequence. Specifically, according to the first part-of-speech sequence, at least one second part-of-speech sequence can be selected from the second part-of-speech sequences respectively corresponding to multiple texts as the tag sequence pattern.
[0060] Exemplarily, the first part-of-speech sequence can be used as the mining object, and the frequent sequence pattern of the first part-of-speech sequence can be mined based on an algorithm. Then, according to the frequent sequence pattern, the second part-of-speech sequences that meet the frequent sequence pattern are selected from the second part-of-speech sequences as the tag sequence pattern.
[0061] Step S240: Determine the tags of the segmented words without annotations in the target text according to the tag sequence pattern, so as to generate the entity relationship extraction result of the target text according to the tags in the target text.
[0062] The target text is the text for which entity relationships are to be extracted. The target text can be any one or any number of texts among the multiple texts. After processing the multiple texts to obtain the tag sequence pattern, the tags of the segmented words without annotations in the target text can be determined by using the tag sequence pattern.
[0063] Exemplarily, the tags in the tag sequence pattern corresponding to the target text can be retrieved by obtaining the tag sequence pattern corresponding to the target text. For example, assuming that the tag sequence pattern corresponding to the target text includes tag sequence pattern A and tag sequence pattern B, then the tags in tag sequence pattern A and the tags in tag sequence pattern B can be retrieved, so as to obtain the tags of the target text. The unlabeled word segments in the target text are labeled with the retrieved tags to obtain new tags in the target text. At the same time, combined with the tags of the labeled word segments in the target text, the entity relationship extraction result of the target text is generated.
[0064] In the entity relationship extraction method of the embodiment of the present application, through word segmentation and part-of-speech tagging processing on multiple texts, the first part-of-speech sequences corresponding to the multiple texts are obtained, and the tags of the labeled word segments in each text are mapped to the part-of-speech elements of the first part-of-speech sequence to generate the second part-of-speech sequence corresponding to each text. According to the part-of-speech elements in the first part-of-speech sequence, at least one second part-of-speech sequence is selected from the second part-of-speech sequences corresponding to each text as the tag sequence pattern. Finally, according to the tag sequence pattern, the tags of the unlabeled word segments in the target text are determined, and combined with the tags of the labeled word segments in the target text, the entity relationship extraction result of the target text is generated, which can achieve efficient and accurate entity relationship extraction, with good flexibility and operability.
[0065] In an embodiment of the present application, the first part-of-speech sequences corresponding to multiple texts can be mined based on the sequence pattern mining method to obtain frequent sequence patterns. Then, according to the frequent sequence patterns, at least one second part-of-speech sequence is selected from the second part-of-speech sequences corresponding to multiple texts as the tag sequence pattern. Refer to Figure 3 , step S230 may specifically include steps S2301 - S2303, which are described in detail as follows:
[0066] Step S2301: Select the part-of-speech elements in the first part-of-speech sequences corresponding to the multiple texts whose first support in the multiple texts is greater than the first threshold to obtain the third part-of-speech sequences corresponding to the multiple texts.
[0067] Sequence pattern mining is to mine all frequent sequences in the sequence database whose support is greater than the minimum support threshold given by the user under the given sequence database and the minimum support threshold given by the user, aiming to discover the frequent sequence patterns in the sequence database.
[0068] Among them, the support of a sequence refers to the support of sequence α in sequence database S, which is the ratio of the number of sequences containing sequence α in sequence database S to the total number of sequences in sequence database S, denoted as support(α). If the support of sequence s is greater than or equal to the minimum support threshold min_sup, then sequence s is called a sequence pattern (frequent sequence). For example, if the minimum support threshold is 2, then a sequence that appears more than twice is considered frequent and is a sequence that needs to be mined.
[0069] Specifically in this step, in order to perform sequence pattern mining on the first part-of-speech sequences respectively corresponding to multiple texts and generate frequent sequence patterns, first, part-of-speech elements with a first support greater than a first threshold in the first part-of-speech sequences respectively corresponding to multiple texts can be selected to obtain third part-of-speech sequences respectively corresponding to multiple texts. The third part-of-speech sequences do not include part-of-speech elements that do not meet the first threshold. In this way, when performing sequence pattern mining on multiple first part-of-speech sequences, not only can frequent sequences with a support greater than the first threshold in multiple first part-of-speech sequences be mined, but also the support of the part-of-speech elements in the frequent sequences meets the requirements of the first threshold.
[0070] In an embodiment of the present application, the support of part-of-speech elements in the first part-of-speech sequences respectively corresponding to multiple texts can be calculated through the text quantity ratio. In this embodiment, as Figure 4 shown, the method further includes step S410-step S420, which are described in detail as follows:
[0071] Step S410: According to the part-of-speech elements in the first part-of-speech sequences respectively corresponding to the multiple texts, count the number of texts containing each part-of-speech element in the multiple texts;
[0072] Step S420: Calculate the ratio between the number of texts containing each part-of-speech element and the total number of the multiple texts to obtain the first support of each part-of-speech element in the multiple texts.
[0073] In this embodiment, for each part-of-speech element, the number of texts containing each part-of-speech element can be counted in multiple texts, and then the ratio between the number of texts containing each part-of-speech element and the total number of multiple texts can be calculated to obtain the first support of each part-of-speech element in multiple texts.
[0074] For example, assume that there are four texts, and the first part-of-speech sequences corresponding to the four texts are " / n / u / n / v / nr", " / n / u / n / n / n / v / nr", " / p / ns / u / n / ns / f / r / n / d / d / v", and " / p / ns / u / n / ns / v / n / a / n" respectively. Then, it can be statistically obtained that the number of texts containing the part-of-speech element n is 4, the number of texts containing the part-of-speech element u is 4, the number of texts containing the part-of-speech element v is 4, the number of texts containing the part-of-speech element nr is 2, the number of texts containing the part-of-speech element ns is 2, the number of texts containing the part-of-speech element p is 2, the number of texts containing the part-of-speech element f is 1, the number of texts containing the part-of-speech element a is 1, the number of texts containing the part-of-speech element d is 1, and the number of texts containing the part-of-speech element r is 1.
[0075] After statistically obtaining the number of texts containing each part-of-speech element, the first support degree of the part-of-speech element n can be calculated as 1, the first support degree of the part-of-speech element u is 1, the first support degree of the part-of-speech element v is 1, the first support degree of the part-of-speech element nr is 1 / 2, the first support degree of the part-of-speech element ns is 1 / 2, the first support degree of the part-of-speech element p is 1 / 2, the first support degree of the part-of-speech element f is 1 / 4, the first support degree of the part-of-speech element a is 1 / 4, the first support degree of the part-of-speech element d is 1 / 4, and the first support degree of the part-of-speech element r is 1 / 4.
[0076] Continue to refer to Figure 3 , in step S2302, sequence pattern mining is performed on the third part-of-speech sequences corresponding to the multiple texts to generate frequent sequence patterns.
[0077] Sequence pattern mining can mine all frequent sequences in the sequence database whose support degree is greater than the minimum support degree threshold. After sequence pattern mining, tens of thousands of sequence patterns will be generated. Therefore, it is necessary to analyze each frequent sequence to discover the frequent sequence patterns in the sequence database.
[0078] In this embodiment, sequence pattern mining can be performed on the third part-of-speech sequences corresponding to the multiple texts based on the sequence pattern mining algorithm. The sequence pattern mining algorithm is not specifically limited in this embodiment of the present application.
[0079] In an embodiment of the present application, sequence pattern mining can be performed on the third part-of-speech sequences corresponding to the multiple texts based on the PrefixSpan algorithm to generate frequent sequence patterns. The PrefixSpan algorithm is a kind of sequence pattern mining algorithm. The following introduces the PrefixSpan algorithm process:
[0080] Input: sequence data set S and support degree threshold α
[0081] Output: All frequent sequence sets that meet the support requirements
[0082] (1) Find all prefixes of length 1 and their corresponding projected databases;
[0083] (2) Count the prefixes of length 1, delete the items corresponding to the prefixes with support lower than the threshold α from the dataset S, and at the same time obtain all frequent 1-item sequences, i = 1;
[0084] (3) Recursively mine for each prefix of length i that meets the support requirements:
[0085] a) Find the projected database corresponding to the prefix. If the projected database is empty, recursively return;
[0086] b) Count the support counts of each item in the corresponding projected database. If the support counts of all items are lower than the threshold α, recursively return;
[0087] c) Combine each single item that meets the support count with the current prefix to obtain several new prefixes;
[0088] d) Let i = i + 1, and the prefixes be the various prefixes after combining single items, and recursively execute step c) respectively.
[0089] Specifically in this embodiment, referring to Figure 5 , step S2302 may specifically include steps S23021 - S23023, which are described in detail as follows:
[0090] Step S23021: Select a part-of-speech element as a prefix from the third part-of-speech sequences respectively corresponding to the multiple texts, and determine at least one suffix corresponding to the prefix. The at least one suffix includes the part-of-speech elements located after the prefix in the third part-of-speech sequence, and the position order of the included part-of-speech elements is the same as the position order in the third part-of-speech sequence;
[0091] Step S23022: Select a part-of-speech element with a second support greater than the first threshold in the at least one suffix and add it to the prefix to obtain a new prefix, and continue to determine a new suffix corresponding to the new prefix until no part-of-speech element with a second support greater than the threshold can be selected from the determined new suffix;
[0092] Step S23023: Generate the frequent sequence pattern according to the obtained multiple prefixes.
[0093] The following uses an example to illustrate steps S23021 - S23022 in this embodiment:
[0094] Suppose the third part-of-speech sequences corresponding to three texts are shown in Table 1 below:
[0095] The third part-of-speech sequences corresponding to the three texts / n / u / n / v / n / u / v / nr / p / ns / d / v
[0096] Table 1
[0097] First, select a part-of-speech element from the third part-of-speech sequences corresponding to the three texts as a prefix, and determine at least one suffix corresponding to it. The at least one suffix contains the part-of-speech elements in the third part-of-speech sequence that are after the prefix, and the order of the contained part-of-speech elements is the same as the order in the third part-of-speech sequence. The results shown in Table 2 can be obtained:
[0098]
[0099]
[0100] Table 2
[0101] Second, select a part-of-speech element in the at least one suffix whose second support degree is greater than the first threshold and add it to the prefix to obtain a new prefix, and continue to determine the new suffix corresponding to the new prefix.
[0102] Taking the prefix " / n" in Table 2 as an example, assuming the first threshold is 0.4, it can be determined that the part-of-speech elements whose second support degree is greater than the first threshold in the suffix corresponding to the prefix " / n" include "u" and "v". Then, "u" and "v" can be added to the prefix " / n" respectively, and the results shown in Table 3 can be obtained:
[0103]
[0104] Table 3
[0105] For the prefix " / n / u" in Table 3, there is a part-of-speech element "v" in its corresponding suffix whose second support degree is greater than the first threshold, so "v" can be added to the prefix " / n / u". For the prefix " / n / v", there is no part-of-speech element in its corresponding suffix whose second support degree is greater than the second threshold. Therefore, the mining of the prefix " / n / v" ends, and the results shown in Table 4 can be obtained:
[0106] Three prefixes Corresponding suffix / n / u / v nr
[0107] Table 4
[0108] For the prefix " / n / u / v" in Table 4, there is no part-of-speech element in its corresponding suffix whose second support degree is greater than the first threshold. Therefore, the mining of the prefix " / n / u / v" ends.
[0109] In one embodiment of the present application, the second support degree of a part-of-speech element in at least one suffix can be calculated by the number of suffixes and the total number of texts. For example, Figure 6 as shown, it includes steps S610 - S620, which are described in detail as follows:
[0110] Step S610: According to the part-of-speech elements in the at least one suffix, count the number of suffixes containing each part-of-speech element in the at least one suffix;
[0111] Step S620: Calculate the ratio between the number of suffixes containing each part-of-speech element and the total number of the multiple texts to obtain the second support degree of each part-of-speech element in the at least one suffix.
[0112] In this embodiment, for each part-of-speech element in the suffix, the number of suffixes containing each part-of-speech element can be counted in at least one suffix, and the ratio between the number of suffixes containing each part-of-speech element and the total number of multiple texts can be calculated to obtain the second support degree of each part-of-speech element in at least one suffix.
[0113] For example, for the prefix " / n" in Table 2 above, its corresponding suffixes are " / u / n / v" and " / u / / v / nr", and the part-of-speech elements in the suffixes are "u", "n", "v", and "nr". Among them, the number of suffixes containing the part-of-speech element "u" is 2, the number of suffixes containing the part-of-speech element "n" is 1, the number of suffixes containing the part-of-speech element "v" is 2, and the number of suffixes containing the part-of-speech element "nr" is 1. After counting the number of suffixes containing each part-of-speech element, the second support degree of the part-of-speech element "u" can be calculated as 2 / 4, the second support degree of the part-of-speech element "n" is 1 / 4, the second support degree of the part-of-speech element "v" is 2 / 4, and the second support degree of the part-of-speech element "nr" is 1 / 4. If the first threshold is 0.4, then the part-of-speech elements with a second support degree greater than the first threshold are "u" and "v".
[0114] Continue to refer to Figure 5 , in step S23023, according to the obtained multiple prefixes, generate the frequent sequence pattern.
[0115] After sequence pattern mining, tens of thousands of prefixes will be generated, and each prefix needs to be analyzed. Among the obtained multiple prefixes, a large number of prefixes are redundant. Therefore, the redundant prefixes can be removed, and the remaining prefixes are used as the frequent sequence pattern.
[0116] In one embodiment, step S23023 may specifically include:
[0117] If there is a target prefix among the multiple prefixes that contains the part-of-speech elements in other prefixes and the position order of the contained part-of-speech elements is the same as that in the other prefixes, then the target prefix is taken as the frequent sequence pattern.
[0118] If all the item sets of a sequence A can be found in the item sets of sequence B, then A is a subsequence of B. According to this definition, for sequence A = {a1, a2,... a n} and sequence B = {b1, b2,... b n}, n ≤ m, if there exists a sequence of numbers 1 ≤ j1 ≤ j2 ≤... ≤ j n ≤ m that satisfies then A is said to be a subsequence of B. Conversely, B is a super-sequence of A.
[0119] Specifically in this embodiment, for the mined prefixes, if there is a target prefix among the multiple prefixes that contains the part-of-speech elements in other prefixes and the position order of the contained part-of-speech elements is the same as that in the other prefixes, then the target prefix is taken as the frequent sequence pattern.
[0120] Continue to refer to Figure 3 In step S2303, according to the frequent sequence pattern, select at least one second part-of-speech sequence from the second part-of-speech sequences corresponding to the multiple texts as the label sequence pattern.
[0121] The frequent sequence pattern is obtained by performing sequence pattern mining on the first part-of-speech sequences corresponding to multiple texts. After obtaining the frequent sequence pattern, at least one second part-of-speech sequence can be further selected from the second part-of-speech sequences corresponding to multiple texts as the label sequence pattern.
[0122] In an embodiment of the present application, as Figure 7 shown, step S2303 may specifically include steps S23031 - S23033, which are described in detail as follows:
[0123] Step S23031: Select at least one target part-of-speech sequence from the second part-of-speech sequences corresponding to the multiple texts. The at least one target part-of-speech sequence contains the part-of-speech elements in the frequent sequence pattern, and the position order of the contained part-of-speech elements is the same as that in the frequent sequence pattern.
[0124] Specifically, select at least one target part-of-speech sequence from the second part-of-speech sequences corresponding to multiple texts. The selected target part-of-speech sequence contains the part-of-speech elements in the frequent sequence pattern, and the position order of the contained part-of-speech elements is the same as that in the frequent sequence pattern.
[0125] Step S23032: Calculate the ratio of the number of tags in each target part-of-speech sequence to the sum of the number of tags in the at least one target part-of-speech sequence, and obtain the confidence corresponding to each target part-of-speech sequence.
[0126] After selecting at least one target part-of-speech sequence from the second part-of-speech sequences corresponding to multiple texts, the ratio of the number of tags in each target part-of-speech sequence to the sum of the number of tags in the at least one target part-of-speech sequence can be calculated, and the calculated ratio is used as the confidence corresponding to each target part-of-speech sequence.
[0127] Step S23033: Use the target part-of-speech sequence with the confidence greater than the second threshold as the tag sequence pattern.
[0128] In this step, select the target part-of-speech sequence with the confidence greater than the second threshold as the tag sequence pattern.
[0129] In the technical solutions of the above embodiments, during the process of mining sequence patterns for the first part-of-speech sequences corresponding to multiple texts, the idea of "snowballing" is used to mine sequence patterns in multiple rounds of iteration, and the first support degree of part-of-speech elements in multiple texts is set. Finally, the accuracy of the mined frequent sequence patterns is ensured. After obtaining the frequent sequence patterns, at least one target part-of-speech sequence can be selected from the second part-of-speech sequences corresponding to multiple texts according to the frequent sequence patterns, and combined with the confidence of the target part-of-speech sequences, and finally the tag sequence pattern is generated, ensuring the reliability of the generated tag sequence pattern.
[0130] In an embodiment of the present application, as Figure 8 shown, calculating the sum of the number of tags in at least one target part-of-speech sequence is obtained by the number of tags with different position serial numbers. In this embodiment, it specifically includes Step S810 - Step S820:
[0131] Step S810: Obtain the position serial numbers corresponding to the tags in each target part-of-speech sequence in each target part-of-speech sequence;
[0132] Step S820: Sum the number of tags with different position serial numbers in the at least one target part-of-speech sequence to obtain the sum of the number of tags in the at least one target part-of-speech sequence.
[0133] The following is an example to illustrate this embodiment:
[0134] Suppose the target part-of-speech sequences include two, namely "# / n / u / n / v# / nr" and "# / n / u* / n / n / n / v / nr" respectively. Among them, the position numbers of the label "#" in the target part-of-speech sequence "# / n / u / n / v# / nr" are 1 and 5, and the position number of the label "#" in the target part-of-speech sequence "# / n / u* / n / n / n / v / nr" is 1, and the position number of the label "*" is 3. Then, the sum of the numbers of labels in the target part-of-speech sequence "# / n / u / n / v# / nr" and the target part-of-speech sequence "# / n / u* / n / n / n / v / nr" can be obtained as 3.
[0135] In an embodiment of the present application, as Figure 9 shown, determining the labels of the undetermined word segmentation in the target text according to the label sequence pattern may specifically include step S2401-step S2402, which are described in detail as follows:
[0136] Step S2401: According to the first part-of-speech sequence corresponding to the target text, select at least one target label sequence pattern from the label sequence patterns, where the at least one target label sequence pattern contains the part-of-speech elements in the first part-of-speech sequence corresponding to the target text, and the position order of the contained part-of-speech elements is the same as the position order in the first part-of-speech sequence corresponding to the target text.
[0137] The target text is the text for which entity relationships are to be extracted. The target text can be any one text or any multiple texts among multiple texts. After obtaining the label sequence pattern by processing multiple texts, the label sequence pattern can be used to determine the labels of the undetermined word segmentation in the target text.
[0138] Specifically, according to the first part-of-speech sequence corresponding to the target text, select at least one target label sequence pattern from the label sequence patterns. The selected at least one target label sequence pattern contains the part-of-speech elements in the first part-of-speech sequence corresponding to the target text, and the position order of the contained part-of-speech elements is the same as the position order in the first part-of-speech sequence corresponding to the target text.
[0139] Step S2402: Determine the labels of the undetermined word segmentation in the target text according to the at least one target label sequence pattern.
[0140] Specifically, the tags of the target text that are not segmented can be determined according to the tags in at least one target tag sequence pattern. For example, there are two target tag sequence patterns, namely "# / n / u / n / v# / nr" and "# / n / u* / n / n / n / v / nr". The tag at the first position of the target tag sequence pattern "# / n / u / n / v# / nr" is "#", and the tag at the fifth position is "#". The tag at the first position of the target tag sequence pattern "# / n / u* / n / n / n / v / nr" is "#", and the tag at the third position is "*". Since the segments at the first position and the fifth position in the target text are already tagged, it can be determined that the tag at the third position in the target text that is not segmented is "*".
[0141] The implementation details of the technical solution of the embodiments of the present application will be described in detail below with four specific texts as examples:
[0142] First, the server obtains the following four texts as shown in Table 5. By performing word segmentation processing and stop word removal processing on the four texts respectively, the following word segmentation sequences as shown in Table 6 can be obtained. Furthermore, by performing part-of-speech tagging processing on each segment in the word segmentation sequence, the following first part-of-speech sequences corresponding to the multiple texts as shown in Table 7 can be obtained.
[0143] Four texts The composer of Song a is Singer A The lyricist of Song a is really Singer B In addition to Shareholder Company b of Company a, other companies have also invested As an investor in Company c, Company d is a large enterprise
[0144] Table 5
[0145] The word segmentation sequences corresponding to the four texts Song a / of / composer / is / Singer A Song a / of / lyricist / person / really / is / Singer B In addition to / Company a / of / shareholder / Company b / outside / other / company / also / have / invested As / Company c / of / investor / Company d / is / a / large / enterprise
[0146] Table 6
[0147] The first part-of-speech sequences corresponding to the four texts / n / u / n / v / nr / n / u / n / n / n / v / nr / p / ns / u / n / ns / f / r / n / d / d / v / p / ns / u / n / ns / v / n / a / n
[0148] Table 7
[0149] Second, for the texts that have been segmented and tagged among the above four texts, there are song a, singer A, company b, company c, lyricist, and shareholder. Among them, song a, singer A, company b, and company c are tagged with entity tags, and lyricist and shareholder are tagged with entity relationship tags. Map the entity tags of song a, singer A, company b, and company c to the part-of-speech elements of the corresponding first part-of-speech sequence, and map the entity relationship tags of lyricist and shareholder to the part-of-speech elements of the corresponding first part-of-speech sequence. The entity tags are represented by "#", and the entity relationship tags are represented by "*". Then, the following second part-of-speech sequence as shown in Table 8 can be obtained.
[0150] The second part-of-speech sequences corresponding to the four texts # / n / u / n / v# / nr # / n / u* / n / n / n / v / nr / p / ns / u* / n# / ns / f / r / n / d / d / v / p# / ns / u / n / ns / v / n / a / n
[0151] Table 8
[0152] Third, calculate the ratio between the number of texts containing the part-of-speech elements in the first part-of-speech sequence and the total number of the four texts to obtain the first support degree of each part-of-speech element in the four texts, as shown in Table 9. Then, select the part-of-speech elements with a first support degree greater than the first threshold, where the first threshold is 0.4, to obtain the third part-of-speech sequences corresponding to the four texts respectively, as shown in Table 10.
[0153] Part-of-speech elements First support degree / n 4 / 4 / u 4 / 4 / v 4 / 4 / nr 2 / 4 / ns 2 / 4 / p 2 / 4 / f 1 / 4 / a 1 / 4 / d 1 / 4 / r 1 / 4
[0154] Table 9
[0155] The third part-of-speech sequences corresponding to the four texts / n / u / n / v / nr / n / u / n / n / n / v / nr / p / ns / u / n / ns / n / v / p / ns / u / n / ns / v / n / n
[0156] Table 10
[0157] Perform sequence pattern mining on the third part-of-speech sequences corresponding to the four texts respectively. Specifically: First, construct a one-item prefix and its corresponding suffix from the part-of-speech elements in the third part-of-speech sequence, and the result is shown in Table 11.
[0158]
[0159]
[0160] Table 11
[0161] Taking the one-item prefixes in Table 11 as " / n" and " / p" as examples, select a part-of-speech element with a second support degree greater than the first threshold in the corresponding suffixes of " / n" and " / p" and add it to " / n" and " / p", and the result shown in Table 12 can be obtained.
[0162]
[0163] Table 12
[0164] Taking the two-item prefixes in Table 12 as " / n / u" and " / p / ns" as examples, select a part-of-speech element with a second support degree greater than the first threshold in the suffixes of " / n / u" and " / p / ns" and add it to " / n / u" and " / p / ns", and the result shown in Table 13 can be obtained.
[0165]
[0166] Table 13
[0167] Taking the three-item prefixes in Table 13 as " / n / u / n" and " / p / ns / u" as examples, select a part-of-speech element with a second support degree greater than the first threshold in the suffixes of " / n / u / n" and " / p / ns / u" and add it to " / n / u / n" and " / p / ns / u", and the result shown in Table 14 can be obtained.
[0168]
[0169] Table 14
[0170] Taking the four prefixes in Table 14, namely, " / n / u / n / v" and " / p / ns / u / ns" as examples, one part-of-speech element that has a second support degree greater than the first threshold in the suffix is selected from " / n / u / n / v" and " / p / ns / u / ns" and added to " / n / u / n / v" and " / p / ns / u / ns", and the results shown in Table 15 can be obtained.
[0171]
[0172] Table 15
[0173] As can be seen from Table 15, only for " / p / ns / u / n / ns" does the part-of-speech element in the corresponding suffix have a second support degree greater than the first threshold. Therefore, the results shown in Table 16 can be further obtained.
[0174]
[0175] Table 16
[0176] Based on the one-item prefix, two-item prefix, three-item prefix, four-item prefix, five-item prefix, and six-item prefix, the frequent sequence patterns shown in Table 17 can be finally obtained.
[0177] Frequent sequence patterns / n / u / n / v / nr / p / ns / u / n / ns / n / p / ns / u / n / ns / v
[0178] Table 17
[0179] Fourth, after obtaining the frequent sequence patterns, at least one target part-of-speech sequence is selected from the second part-of-speech sequences corresponding to the four texts shown in Table 4. The target part-of-speech sequence contains the part-of-speech elements in the frequent sequence pattern, and the order of the contained part-of-speech elements is the same as the order in the frequent sequence pattern. For example, for the frequent sequence pattern " / n / u / n / v / nr", the target part-of-speech sequences can be determined as "# / n / u / n / v# / nr" and "# / n / u* / n / n / n / v / nr".
[0180] Furthermore, the confidence corresponding to the target part-of-speech sequence can be calculated. If the confidence is greater than the second threshold, the target part-of-speech sequence can be used as the label sequence pattern. For example, for the frequent sequence pattern " / n / u / n / v / nr", the confidence of the selected target part-of-speech sequence "# / n / u / n / v# / nr" is 2 / 3, and the confidence of the selected target part-of-speech sequence "# / n / u* / n / n / n / v / nr" is 2 / 3. If the second threshold is 1 / 3, the target part-of-speech sequences "# / n / u / n / v# / nr" and "# / n / u* / n / n / n / v / nr" can be used as the label sequence patterns.
[0181] Through the above method, the label sequence patterns are finally obtained as shown in Table 18.
[0182] Label sequence patterns # / n / u / n / v# / nr # / n / u* / n / n / n / v / nr / p / ns / u* / n# / ns / f / r / n / d / d / v / p# / ns / u / n / ns / v / n / a / n
[0183] Table 18
[0184] Fifth, determine the labels of the unannotated word segments in the target text according to the label sequence patterns, so as to generate the entity relationship extraction result of the target text based on the labels in the target text.
[0185] The target text can be the four texts shown in Table 1. Select at least one target label sequence pattern from the label sequence patterns. At least one target label sequence pattern contains the part-of-speech elements in the first part-of-speech sequence corresponding to the target text, and the position order of the contained part-of-speech elements is the same as the position order in the first part-of-speech sequence corresponding to the target text.
[0186] Obtain the labels of the unannotated word segments in the target text according to the labels of at least one target label sequence pattern, so as to generate the entity relationship extraction result of the target text.
[0187] For example, for the target text "The composer of song a is singer A", the first part-of-speech sequence corresponding to this target text is " / n / u / n / v / nr", and the target label sequence patterns selected according to this first part-of-speech sequence are "# / n / u / n / v# / nr" and "# / n / u* / n / n / n / v / nr". Since the word segments at the first position and the fifth position in this target text have been labeled, and according to the target label sequence patterns, the word segment at the third position in this target text that has not been labeled can also be labeled with "*", so that the label of "composer" can be the entity relationship label.
[0188] Through the above process, it is possible to label the unlabeled word segmentation tags in the four texts, label the entity relationship tags for "composing music", label the entity tags for "Singer B", label the entity tags for "Company a", label the entity relationship tags for "investor", and label the entity tags for "Company d". Combining with the tags of the already labeled word segmentation, the entity relationship extraction results of the four texts can be obtained as follows
[0189] as shown in Table 19.
[0190] Relationship entities Relationship Relationship entities Song a Compose Singer A Song a Lyricize Singer B Company a Shareholder Company b Company c Investor Company d
[0191] Table 19
[0192] The following introduces the device embodiments of the present application, which can be used to execute the entity relationship extraction method in the above embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the embodiments of the entity relationship extraction method above of the present application.
[0193] Figure 10 shows a block diagram of an entity relationship extraction device according to an embodiment of the present application. Refer to Figure 10 As shown, an entity relationship extraction device 1000 according to an embodiment of the present application includes: an acquisition unit 1002, a generation unit 1004, a selection unit 1006, and a determination unit 1008.
[0194] Among them, the acquisition unit 1002 is configured to acquire the first part-of-speech sequences corresponding to multiple texts respectively, and the first part-of-speech sequence corresponding to each text includes the part-of-speech elements corresponding to each word segmentation in the word segmentation result of each text; the generation unit 1004 is configured to map the tags of the already labeled word segmentation in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text to generate the second part-of-speech sequence corresponding to each text, and the tags include entity tags and entity relationship tags; the selection unit 1006 is configured to select at least one second part-of-speech sequence as the tag sequence pattern from the second part-of-speech sequences corresponding to the multiple texts respectively according to the first part-of-speech sequences corresponding to the multiple texts respectively; the determination unit 1008 is configured to determine the tags of the unlabeled word segmentation in the target text according to the tag sequence pattern, so as to generate the entity relationship extraction result of the target text according to the tags in the target text.
[0195] In some embodiments of the present application, the selection unit 1006 includes: a first selection subunit configured to select, from the first part-of-speech sequences respectively corresponding to the multiple texts, part-of-speech elements with a first support degree greater than a first threshold in the multiple texts, so as to obtain third part-of-speech sequences respectively corresponding to the multiple texts; a mining subunit configured to perform sequence pattern mining on the third part-of-speech sequences respectively corresponding to the multiple texts to generate frequent sequence patterns; and a second selection subunit configured to select at least one second part-of-speech sequence as the tag sequence pattern from the second part-of-speech sequences respectively corresponding to the multiple texts according to the frequent sequence patterns.
[0196] In some embodiments of the present application, the first selection subunit is further configured to: count the number of texts containing each part-of-speech element in the multiple texts according to the part-of-speech elements in the first part-of-speech sequences respectively corresponding to the multiple texts; and calculate a ratio between the number of texts containing each part-of-speech element and the total number of the multiple texts to obtain the first support degree of each part-of-speech element in the multiple texts.
[0197] In some embodiments of the present application, the mining subunit is further configured to: select a part-of-speech element from the third part-of-speech sequences respectively corresponding to the multiple texts as a prefix, and determine at least one suffix corresponding to the prefix, where the at least one suffix includes part-of-speech elements located after the prefix in the third part-of-speech sequences, and the position order of the included part-of-speech elements is the same as the position order in the third part-of-speech sequences; select a part-of-speech element with a second support degree greater than the first threshold in the at least one suffix and add it to the prefix to obtain a new prefix, and continue to determine a new suffix corresponding to the new prefix until no part-of-speech element with a second support degree greater than the threshold can be selected from the determined new suffixes; and generate the frequent sequence patterns according to the obtained multiple prefixes.
[0198] In some embodiments of the present application, the mining subunit is further configured to: if there is a target prefix in the multiple prefixes that includes part-of-speech elements in other prefixes and the position order of the included part-of-speech elements is the same as the position order in the other prefixes, use the target prefix as the frequent sequence pattern.
[0199] In some embodiments of the present application, the mining subunit is further configured to: count the number of suffixes containing each part-of-speech element in the at least one suffix according to the part-of-speech elements in the at least one suffix; and calculate a ratio between the number of suffixes containing each part-of-speech element and the total number of the multiple texts to obtain the second support degree of each part-of-speech element in the at least one suffix.
[0200] In some embodiments of the present application, the second selection subunit is further configured to: select at least one target part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the multiple texts according to the frequent sequence pattern, where the at least one target part-of-speech sequence includes the part-of-speech elements in the frequent sequence pattern, and the position order of the included part-of-speech elements is the same as the position order in the frequent sequence pattern; calculate the ratio of the number of tags in each target part-of-speech sequence to the sum of the number of tags in the at least one target part-of-speech sequence to obtain the confidence corresponding to each target part-of-speech sequence; and use the target part-of-speech sequence with a confidence greater than a second threshold as the tag sequence pattern.
[0201] In some embodiments of the present application, the second selection subunit is further configured to: obtain the position serial numbers corresponding to the tags in each target part-of-speech sequence in each target part-of-speech sequence; and sum the number of tags with different position serial numbers in the at least one target part-of-speech sequence to obtain the sum of the number of tags in the at least one target part-of-speech sequence.
[0202] In some embodiments of the present application, the determination unit 1008 is further configured to: select at least one target tag sequence pattern from the tag sequence pattern according to the first part-of-speech sequence corresponding to the target text, where the at least one target tag sequence pattern includes the part-of-speech elements in the first part-of-speech sequence corresponding to the target text, and the position order of the included part-of-speech elements is the same as the position order in the first part-of-speech sequence corresponding to the target text; and determine the tags of the target text that are not segmented and labeled according to the at least one target tag sequence pattern.
[0203] Figure 11 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing embodiments of the present application.
[0204] It should be noted that Figure 11 The computer system 1100 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0205] Such as Figure 11As shown, computer system 1100 includes a Central Processing Unit (CPU) 1101, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 1102 or programs loaded from a storage section 1108 into a Random Access Memory (RAM) 1103, such as executing the methods described in the above embodiments. In the RAM 1103, various programs and data required for system operation are also stored. The CPU 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An Input / Output (I / O) interface 1105 is also connected to the bus 1104.
[0206] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage section 1108 as needed.
[0207] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by a Central Processing Unit (CPU) 1101, various functions defined in the system of the present application are executed.
[0208] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable computer program is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0210] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0211] As another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the above embodiments.
[0212] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0213] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of this application.
[0214] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include known common general knowledge or conventional technical means in the technical field not disclosed in this application.
[0215] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.
Claims
1. An entity relationship extraction method, characterized in that, The method includes: Obtaining first part-of-speech sequences corresponding to multiple texts, where the first part-of-speech sequence corresponding to each text includes part-of-speech elements corresponding to each word segment in the word segmentation result of each text; Mapping the tags of the word segments with annotations in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text, generating a second part-of-speech sequence corresponding to each text, where the tags include entity tags and entity relationship tags; Generating frequent sequence patterns according to the first part-of-speech sequences corresponding to the multiple texts, and selecting at least one target part-of-speech sequence from the second part-of-speech sequences corresponding to the multiple texts according to the frequent sequence patterns, where the at least one target part-of-speech sequence includes the part-of-speech elements in the frequent sequence patterns, and the position order of the included part-of-speech elements is the same as the position order in the frequent sequence patterns; Obtaining the position serial numbers corresponding to the tags in each target part-of-speech sequence in each target part-of-speech sequence, and summing the number of tags with different position serial numbers in the at least one target part-of-speech sequence to obtain the sum of the number of tags in the at least one target part-of-speech sequence; Calculating the ratio of the number of tags in each target part-of-speech sequence to the sum of the number of tags in the at least one target part-of-speech sequence to obtain the confidence corresponding to each target part-of-speech sequence, and taking the target part-of-speech sequence with the confidence greater than the second threshold as the tag sequence pattern; Determining the tags of the word segments without annotations in the target text according to the tag sequence pattern, so as to generate the entity relationship extraction result of the target text according to the tags in the target text.
2. The method according to claim 1, wherein Generating frequent sequence patterns according to the first part-of-speech sequences corresponding to the multiple texts, including: Selecting part-of-speech elements with a first support degree greater than the first threshold in the multiple texts from the first part-of-speech sequences corresponding to the multiple texts to obtain third part-of-speech sequences corresponding to the multiple texts; Performing sequence pattern mining on the third part-of-speech sequences corresponding to the multiple texts to generate the frequent sequence patterns.
3. The method according to claim 2, characterized in that, The method further includes: Counting the number of texts containing each part-of-speech element in the multiple texts according to the part-of-speech elements in the first part-of-speech sequences corresponding to the multiple texts; Calculating the ratio between the number of texts containing each part-of-speech element and the total number of the multiple texts to obtain the first support degree of each part-of-speech element in the multiple texts.
4. The method according to claim 2, wherein Performing sequence pattern mining on the third part-of-speech sequences corresponding to the multiple texts to generate the frequent sequence patterns, including: Selecting a part-of-speech element as a prefix from the third part-of-speech sequences corresponding to the multiple texts, and determining at least one suffix corresponding to the prefix, where the at least one suffix includes the part-of-speech elements in the third part-of-speech sequences located after the prefix, and the position order of the included part-of-speech elements is the same as the position order in the third part-of-speech sequences; Select a part-of-speech element in the at least one suffix whose second support in the at least one suffix is greater than the first threshold, add it to the prefix to obtain a new prefix, and continue to determine a new suffix corresponding to the new prefix until no part-of-speech element with a second support greater than the threshold can be selected from the determined new suffixes; Generate the frequent sequence pattern according to the obtained multiple prefixes.
5. The method according to claim 4, wherein Generating the frequent sequence pattern according to the obtained multiple prefixes includes: If there is a target prefix among the multiple prefixes that contains part-of-speech elements in other prefixes and the position order of the contained part-of-speech elements is the same as that in the other prefixes, then use the target prefix as the frequent sequence pattern.
6. The method according to claim 4, characterized in that The method includes: According to the part-of-speech elements in the at least one suffix, count the number of suffixes containing each part-of-speech element in the at least one suffix; Calculate the ratio between the number of suffixes containing each part-of-speech element and the total number of the multiple texts to obtain the second support of each part-of-speech element in the at least one suffix.
7. The method according to claim 1, wherein Determining the unlabeled word segmentation tags in the target text according to the tag sequence pattern includes: According to the first part-of-speech sequence corresponding to the target text, select at least one target tag sequence pattern from the tag sequence pattern, where the at least one target tag sequence pattern contains the part-of-speech elements in the first part-of-speech sequence corresponding to the target text and the position order of the contained part-of-speech elements is the same as that in the first part-of-speech sequence corresponding to the target text; Determine the unlabeled word segmentation tags in the target text according to the at least one target tag sequence pattern.
8. An entity relationship extraction device, characterized in that, The device includes: An acquisition unit configured to acquire the first part-of-speech sequences respectively corresponding to multiple texts, and each first part-of-speech sequence corresponding to a text contains the part-of-speech elements corresponding to each word segmentation in the word segmentation result of each text; A generation unit configured to map the labeled word segmentation tags in each text to the part-of-speech elements of the first part-of-speech sequence corresponding to each text to generate the second part-of-speech sequence corresponding to each text, and the tags include entity tags and entity relationship tags; A selection unit configured to generate a frequent sequence pattern according to the first part-of-speech sequences respectively corresponding to the multiple texts, and select at least one target part-of-speech sequence from the second part-of-speech sequences respectively corresponding to the multiple texts according to the frequent sequence pattern, where the at least one target part-of-speech sequence contains the part-of-speech elements in the frequent sequence pattern and the position order of the contained part-of-speech elements is the same as that in the frequent sequence pattern; Obtain the position serial numbers corresponding to the tags in each target part-of-speech sequence in each target part-of-speech sequence, and sum the number of tags with different position serial numbers in the at least one target part-of-speech sequence to obtain the sum of the number of tags in the at least one target part-of-speech sequence; Calculate the ratio of the number of tags in each target part-of-speech sequence to the sum of the number of tags in the at least one target part-of-speech sequence to obtain the confidence corresponding to each target part-of-speech sequence, and use the target part-of-speech sequence with the confidence greater than the second threshold as the tag sequence pattern; A determination unit, configured to determine the tags of the unlabeled word segments in the target text according to the tag sequence pattern, so as to generate an entity relationship extraction result of the target text based on the tags in the target text.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 7 by executing the executable instructions.
11. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and a processor of a computer device reads and executes the computer program from the computer-readable storage medium, so that the computer device executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method, device and equipment for carrying out event extraction on text and computer storage medium
CN109815481A
Data identification method and device, storage medium and electronic equipment
CN111368555A