Model training method, entity relationship extraction method and related equipment

By performing feature extraction and comprehensive calculation of loss functions on text, the accuracy of entity relationship extraction at the document level is improved, the problem of low accuracy in existing technologies is solved, and more accurate entity relationship extraction is achieved.

CN120688491APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411639646.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing document-level entity relationship extraction accuracy is not high.

Method used

By extracting features from the text, we obtain the feature representations of words and sentences. We extract entity relationships based on the feature representations of words, determine the probability distribution, and combine the feature representations of entities and sentences to determine the loss, and then train the model.

Benefits of technology

The accuracy of entity relationship extraction at the document level is improved, the introduction of noise is avoided, and the effect of entity relationship extraction is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688491A_ABST
    Figure CN120688491A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an entity relationship extraction method and related equipment. The training method comprises the following steps: performing feature extraction on a first text to obtain a feature representation of each lexical element in the first text and a feature representation of each sentence in the first text; performing entity relationship extraction based on the feature representation of each lexical element to obtain probability distribution of each lexical element position in the output sequence; determining an entity pair in the first text based on the probability distribution of each lexical element position; determining loss based on the probability distribution of each lexical element position, the feature representation of each entity of the entity pair and the feature representation of each sentence; and training a first model based on the loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a model training method, an entity relationship extraction method, and related equipment. Background Art

[0002] Entity relationship extraction is a basic and important task in the field of natural language processing, which refers to the process of automatically identifying the relationship between entities from text.

[0003] Entity relationship extraction initially focused on sentence-level entity relationship extraction. However, sentence-level entity relationship extraction can only process entity relationship extraction within sentences, while the relationship between entities may require cross-sentence reasoning to derive the relationship. Therefore, the industry began to explore document-level entity relationship extraction (also known as paragraph-level entity relationship extraction).

[0004] At present, the implementation algorithms for document-level entity relationship extraction include: graph neural network (GNN)-based methods, joint extraction methods, etc. However, the above methods make the accuracy of document-level entity relationship extraction low. Summary of the Invention

[0005] This application provides a model training method, entity relationship extraction method and related equipment to improve the accuracy of document-level entity relationship extraction.

[0006] In a first aspect, the present application provides a model training method, the method comprising:

[0007] Performing feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text;

[0008] Entity relations are extracted based on the feature representation of each word, and the probability distribution of each word position is obtained;

[0009] Determining entity pairs in the first text based on the probability distribution of each word position;

[0010] Determine the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence;

[0011] The first model is trained based on the loss.

[0012] In a second aspect, the present application provides an entity relationship extraction method, the method comprising:

[0013] Get the second text;

[0014] The trained first model is used to extract entity relationships from the second text to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair, wherein the trained first model is obtained by training using the model training method provided by the first aspect.

[0015] In a third aspect, the present application provides a model training device, the device comprising: a first acquisition unit and a first processing unit;

[0016] A first acquiring unit, configured to acquire a first text;

[0017] A first processing unit is configured to perform feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text;

[0018] The first processing unit is further configured to extract entity relationships based on the feature representation of each word element, and obtain a probability distribution of each word element position;

[0019] The first processing unit is further configured to determine entity pairs in the first text based on the probability distribution of each word unit position;

[0020] The first processing unit is further configured to determine a loss based on a probability distribution of each word position, a feature representation of each entity of the entity pair, and a feature representation of each sentence;

[0021] The first processing unit is further configured to train the first model based on the loss.

[0022] In a fourth aspect, the present application provides an entity relationship extraction device, the device comprising: a second acquisition unit and a second processing unit;

[0023] The second acquiring unit is configured to acquire a second text;

[0024] The second processing unit is used to extract entity relationships from the second text using the trained first model to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair, wherein the trained first model is obtained by training using the model training method provided by the first aspect.

[0025] In a fifth aspect, the present application provides an electronic device comprising: a processor and a memory, the processor being connected to the memory, the memory being used to store computer programs, and the processor being used to execute the computer programs stored in the memory, so that the electronic device performs the method of the first aspect or the second aspect.

[0026] In a sixth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method of the first aspect or the second aspect is performed.

[0027] In a seventh aspect, the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, performs the method of the first aspect or the second aspect.

[0028] The implementation of this application has the following beneficial effects:

[0029] The present application trains a model, specifically: extracting features from the first text to obtain feature representations of each word in the first text and feature representations of each sentence in the first text; then extracting entity relationships based on the feature representations of each word, and obtaining a probability distribution of each word position in the output sequence, that is, based on the entity relationship extraction based on each word feature representation, each word position in the output sequence obtained includes entity pairs and word positions corresponding to the relationships between entities in the entity pairs, then the entity pairs in the first text can be determined based on the probability distribution of each word position; then the loss is determined based on the probability distribution of each word position, the feature representations of each entity in the entity pairs, and the feature representations of each sentence in the first text; and then the first model is trained based on the loss, that is, since each word position in the output sequence obtained includes entity pairs and relationships between entities in the entity pairs, The word position corresponding to the relationship, that is, the probability distribution of the word position corresponding to the entity pair and the probability distribution of the word position corresponding to the relationship between the entities in the entity pair. After obtaining the probability distribution of each word position, the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence in the first text are combined to comprehensively calculate the loss. Compared with directly determining the entity pair and the corresponding entity relationship based on the probability distribution of each word position and then determining the loss, the feature representation of each entity in the entity pair and the feature representation of each sentence in the first text are added to determine the loss. By combining the feature representation of each entity in the entity pair and the feature representation of each sentence, the model can learn multi-dimensional considerations and enhance the evidence for the extraction of entity relationships, thereby improving the accuracy of entity relationship extraction when the trained first model is used for entity relationship extraction.

[0030] In addition, the present application also extracts features from the second text to obtain feature representations of each word element in the second text and feature representations of each sentence in the second text; then, entity relationship extraction is performed based on the feature representation of each word element to obtain a second entity pair in the second text and the relationship between the two entities of the second entity pair; then, based on the feature representation of each sentence in the second text and the feature representation of each entity of the second entity pair, the first sentence corresponding to the second entity pair is determined, wherein the first sentence is a sentence in the second text used to prove the relationship between the two entities of the second entity pair; then, entity relationship extraction is performed based on the feature representation of the first sentence to obtain a third entity pair in the second text and the relationship between the two entities of the third entity pair; finally, based on the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, the third entity pair, the third entity pair The relationship between the two entities of the entity pair is used to determine the entity relationship extraction result corresponding to the second text. That is to say, after performing entity relationship extraction based on the feature representation of each word unit to obtain the second entity pair in the second text and the relationship between the two entities of the second entity pair, the feature representation of each sentence and the feature representation of each entity of the second entity pair are combined to determine the first sentence corresponding to the second entity pair. This can avoid the introduction of noise and improve the accuracy of determining the first sentence. Then, based on the first sentence, the third entity pair and the relationship between the two entities of the third entity pair are predicted, which can enhance the effect of entity relationship extraction. Finally, the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, and the relationship between the two entities of the third entity pair are combined to determine the final entity relationship extraction result, thereby improving the accuracy of entity relationship extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0032] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;

[0033] Figure 2 A training diagram of a model provided in an embodiment of the present application;

[0034] Figure 3 A flowchart of an entity relationship extraction method provided in an embodiment of the present application;

[0035] Figure 4 A schematic diagram of entity relationship extraction provided in an embodiment of the present application;

[0036] Figure 5 A schematic diagram of an entity relationship extraction system provided in an embodiment of the present application;

[0037] Figure 6 A schematic diagram of a model training system provided in an embodiment of the present application;

[0038] Figure 7 A block diagram of the functional units of a model training device provided in an embodiment of the present application;

[0039] Figure 8 A block diagram of the functional units of an entity relationship extraction device provided in an embodiment of the present application;

[0040] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0042] The terms "first," "second," "third," and "fourth," etc., in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a particular order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0043] References herein to "embodiments" mean that a particular feature, result, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0044] From the above background technology, it can be seen that the accuracy of existing document-level entity relationship extraction is not high. Therefore, this application trains a model, specifically: feature extraction is performed on the first text to obtain the feature representation of each word in the first text and the feature representation of each sentence in the first text; then entity relationship extraction is performed based on the feature representation of each word, and the probability distribution of each word position in the output sequence can be obtained, that is, based on the feature representation of each word, each word position in the output sequence obtained by performing entity relationship extraction includes entity pairs and word positions corresponding to the relationship between entities in the entity pairs, then the entity pairs in the first text can be determined based on the probability distribution of each word position; then the loss is determined based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence in the first text; and then the first model is trained based on the loss, that is, since the position of each word in the output sequence obtained is The word positions corresponding to the entity pairs and the relationships between the entities in the entity pairs, that is, the probability distribution of the word positions corresponding to the entity pairs and the probability distribution of the word positions corresponding to the relationships between the entities in the entity pairs, after obtaining the probability distribution of each word position, the loss is comprehensively calculated by combining the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence in the first text. Compared with directly determining the entity pairs and the corresponding entity relationships based on the probability distribution of each word position and then determining the loss, the feature representation of each entity in the entity pair and the feature representation of each sentence in the first text are added to determine the loss. By combining the feature representation of each entity in the entity pair and the feature representation of each sentence, the model can learn multi-dimensional considerations and perform evidence enhancement on the extraction of entity relationships, thereby improving the accuracy of entity relationship extraction when the trained first model is used for entity relationship extraction.

[0045] In addition, the present application also extracts features from the second text to obtain feature representations of each word element in the second text and feature representations of each sentence in the second text; then, entity relationship extraction is performed based on the feature representation of each word element to obtain a second entity pair in the second text and the relationship between the two entities of the second entity pair; then, based on the feature representation of each sentence in the second text and the feature representation of each entity of the second entity pair, the first sentence corresponding to the second entity pair is determined, wherein the first sentence is a sentence in the second text used to prove the relationship between the two entities of the second entity pair; then, entity relationship extraction is performed based on the feature representation of the first sentence to obtain a third entity pair in the second text and the relationship between the two entities of the third entity pair; finally, based on the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, the third entity pair, the third entity pair The relationship between the two entities of the entity pair is used to determine the entity relationship extraction result corresponding to the second text. That is to say, after performing entity relationship extraction based on the feature representation of each word unit to obtain the second entity pair in the second text and the relationship between the two entities of the second entity pair, the feature representation of each sentence and the feature representation of each entity of the second entity pair are combined to determine the first sentence corresponding to the second entity pair. This can avoid the introduction of noise and improve the accuracy of determining the first sentence. Then, based on the first sentence, the third entity pair and the relationship between the two entities of the third entity pair are predicted, which can enhance the effect of entity relationship extraction. Finally, the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, and the relationship between the two entities of the third entity pair are combined to determine the final entity relationship extraction result, thereby improving the accuracy of entity relationship extraction.

[0046] See Figure 1 , Figure 1 A flow chart of a model training method provided in an embodiment of the present application. The method is applied to a model training device, and the method includes but is not limited to steps S101-S105:

[0047] S101 : Perform feature extraction on a first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text.

[0048] In the embodiment of the present application, the number of characters of the first text is greater than the threshold value. The first text may be a document-level text, and the present application does not specifically limit the number of the first text. The present application mainly uses a first text as an example for explanation; the first text includes A sentences and B words The first text (also called input sequence, denoted as D) can be input into an encoder (such as a pre-trained BART model, Bert model, Transformer model, etc., which is not limited in this application), and the feature representation or embedding representation (Embedding) of each word in the first text is output, as shown in formula (1):

[0049] H=[h1,…,h B ]=Encoder([t1,…,t B ])(1)

[0050] Then, based on the feature representation of each word in the first text, the feature representation of each word in each sentence can be determined. Then, based on the feature representation of each word in each sentence, the feature representation of each sentence can be obtained. For example, the LogSumExp method can be used to pool each word in each sentence, and the nth sentence S n As an example, as shown in formula (2):

[0051]

[0052] Among them, token i For sentence S n It should be noted that the method of performing feature extraction on the first text to obtain the feature representation of each word in the first text and the feature representation of each sentence can also adopt other methods, which are not specifically limited in this application.

[0053] S102: Entity relationship extraction is performed based on the feature representation of each word, and a probability distribution of each word position in the output sequence is obtained.

[0054] In an embodiment of the present application, the feature representation of each word unit can be input into a decoder (such as an autoregressive decoder), which is decoded step by step by the decoder to obtain the probability distribution of each word unit position in the output sequence (denoted as X, with a length of T). This will not be elaborated here, and please refer to the corresponding explanation of the embodiment below for details; optionally, after obtaining the feature representation of each word unit, the feature representation of each word unit can also be processed by an attention mechanism to obtain a new feature representation of each word unit, and then the new feature representation of each word unit is input into an autoregressive decoder, which is decoded step by step by the autoregressive decoder to obtain the probability distribution of each word unit position in the output sequence, wherein the probability distribution of each word unit position in the output sequence can represent the probability of each word unit position in the preset vocabulary. The probability distribution under the word, or in other words, the probability distribution of each word position can represent the probability distribution of each word position under each word position in the input sequence D, and the probability distribution of each word position is determined based on the (new) feature representation of each word in the input sequence D and the label corresponding to the word position before each word position. For example, assuming the i-th word position in the output sequence, after the decoder predicts the i-1-th word position, the label corresponding to the i-1-th word position and the (new) feature representation of each word in the input sequence D are used as input for predicting the probability distribution corresponding to the i-th word position, and then the probability distribution corresponding to the i-th word position is output. The specific principles are not elaborated here. Please see the corresponding explanation of the embodiment below.

[0055] S103: Determine entity pairs in the first text based on the probability distribution of each word position.

[0056] In an embodiment of the present application, after obtaining the probability distribution corresponding to each word position in the output sequence, the predicted word corresponding to each word position in the output sequence can be determined based on the probability distribution corresponding to each word position, that is, the output sequence is obtained. For example, the word with the largest probability in the probability distribution corresponding to each word position is determined as the predicted word corresponding to each word position; since the autoregressive decoder extracts entity relationships based on the feature representation of each word in the input sequence, the obtained output sequence includes entity pairs and entity relationships between the two entities of the entity pair, and then the entity pairs in the first text can be determined based on the output sequence, wherein the number of entity pairs can be one or more, and the entity relationship between the two entities of each entity pair can be one or more, which is not limited in this application. The embodiment of the present application is mainly explained by taking one entity pair as an example.

[0057] In an optional embodiment, in order to avoid the presence of entities or entity relationships not mentioned in the input sequence when the first model generates an output sequence, the above-mentioned vocabulary can also be restricted, that is, the words in the vocabulary are restricted to some special tags, which are used for entity and entity relationship extraction modeling. For example, these special tags can be used to indicate entities, entity relationships, and can include entities and entity relationships in the input sequence. That is, at this time, each special tag in the vocabulary corresponds to an entity or entity relationship, and then the output sequence can be generated based on the probability distribution of each word position in the output sequence under the restricted vocabulary. It will not be repeated here; by restricting the vocabulary, it helps the first model to better learn entity relationship extraction, avoid the output sequence generated by the first model including entities or entity relationships that do not appear in the input sequence, and because entity relationship extraction needs to process a large number of entities and entity relationships, by restricting the vocabulary, the computational complexity of the first model can be reduced and the training efficiency of the model can be improved.

[0058] In addition, the position code of the entity in the input sequence can also be determined, and the position code corresponding to the entity in the input sequence is added to the above vocabulary, that is, the vocabulary at this time includes the entity and the position code corresponding to the entity. Then, when the autoregressive decoder is decoding, if the token with the highest probability in the probability distribution corresponding to each word position in the output sequence has a corresponding position code, then based on the position code corresponding to the token with the highest probability, the entity token corresponding to the position code is copied from the input sequence to each word position in the output sequence. The model does not need to generate it again, but directly copies it, which is equivalent to effectively expanding the vocabulary after the above restrictions, which helps the model to correctly capture the relationship between entities, thereby improving the performance of the entity relationship extraction task.

[0059] S104. Determine the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence.

[0060] In an embodiment of the present application, the first loss may be determined based on the probability distribution of each word unit position and the label corresponding to each word unit position. For example:

[0061] First, based on the probability distribution of each word position, determine the conditional probability of each word position under the corresponding label. Since the probability distribution of each word position is based on the (new) feature representation of each word in the input sequence D and the label corresponding to the word position before each word position, the conditional probability of each word position under the corresponding label can also be understood as being based on the (new) feature representation of each word in the input sequence D and the label corresponding to the word position before each word position. For example, for the output sequence X1,...,X1 of length T, T , after determining the X1 to X T-1 After the corresponding probability distribution, we can T-1 The corresponding probability distribution determines X1 to X T-1 The conditional probability under the corresponding label, and then X T The conditional probability under the corresponding label can be expressed as P(X T |D,X1,……,X T-1 ) or expressed as P(X T |D,X<T).

[0062] Then, based on the conditional probability of each word position under the corresponding label, the first loss is determined. For example, the first loss can be determined based on the cross entropy loss function and the conditional probability of each word position under the corresponding label. Optionally, the first loss can also be determined based on the negative log-likelihood function and the conditional probability of each word position under the corresponding label. The first loss is shown in formula (3):

[0063]

[0064] Among them, Loss1(ω) represents the first loss, P(X i |D,X<i) corresponds to the conditional probability of the i-th word position in the output sequence under the corresponding label, X is the output sequence, D is the input sequence, T is the length of the output sequence, X<i corresponds to the label of each word position before the i-th word position in the output sequence X, log() is the logarithmic function, and ω represents the model parameters; that is, by minimizing the negative log-likelihood function, the probability of each word position in the output sequence under the corresponding label is the highest under the model parameters after the model converges, and the corresponding first loss is the smallest, which means that the model has the best fit on the training data. Of course, the first loss can also be calculated using other loss functions, such as the mean square error loss function, the relative entropy loss function, etc., which are not specifically limited in this application.

[0065] Of course, the joint conditional probability corresponding to the output sequence X can also be determined based on the conditional probability of each word position in the output sequence under the corresponding label. For example, the joint conditional probability corresponding to the output sequence X can be determined based on the likelihood function, as shown in formula (4):

[0066]

[0067] Among them, P(X i |D,X<i) corresponds to the conditional probability of the i-th word position in the output sequence under the corresponding label, X is the output sequence, D is the input sequence, T is the length of the output sequence, X<i corresponds to the X in the output sequence X i The label of each word position before; then the first loss can be determined based on the log-likelihood function and the joint conditional probability corresponding to the output sequence X, that is, the model is learned by maximizing the log-likelihood function. For example, the first loss is shown in formula (5):

[0068]

[0069] Among them, Loss1(ω) represents the first loss, P(X i |D,X<i) corresponds to the conditional probability of the i-th word position in the output sequence under the corresponding label, X is the output sequence, D is the input sequence, T is the length of the output sequence, X<i corresponds to the label of each word position before the i-th word position in the output sequence X, log() is the logarithmic function, and ω represents the model parameters; that is, by maximizing the log-likelihood function, the probability of each word position in the output sequence under the corresponding label is the highest under the model parameters after model convergence, and the corresponding first loss is minimized, which means that the model fits the training data best.

[0070] Of course, optionally, the performance results of the first model can also be verified based on the joint conditional probability corresponding to the output sequence X determined above. That is, after the subsequent loss-based model training converges, the performance results of the first model can be verified based on the joint conditional probability corresponding to the output sequence X. For example, if the joint conditional probability corresponding to the output sequence X determined after the first model converges is greater than the first threshold, it means that the model parameters obtained by training at this time meet the experimental requirements, providing data support for technical personnel.

[0071] Then, based on the feature representation of each entity of the entity pair, the feature representation of each sentence, and the pseudo-label of the entity pair, a second loss is determined, wherein the pseudo-label is used to indicate the annotation of the entity of the entity pair for each sentence in the first text. For example, when there is a corresponding entity in the sentence, the pseudo-label is 1, and when there is no corresponding entity in the sentence, the pseudo-label is 0. Specifically, the sentence where each entity of the entity pair is located in the first text can be obtained, and then based on each sentence in the first text and the sentence where each entity of the entity pair is located in the first text, the pseudo-label of the entity pair is generated. Specifically, based on the sentence where each entity of the entity pair in the first text is located, each sentence in the first text is one-hot encoded (One-Hot Encoding) to obtain the pseudo-label of the entity pair, wherein the sentence where each entity of the entity pair is located in the first text is encoded as 1, and the remaining sentences in the first text except the sentence where each entity of the entity pair is located are encoded as 0, that is, the pseudo-label can be regarded as a set of label sequences including 0 and 1. For example, if the first text includes 6 sentences, and the two entities of the entity pair are located in the 3rd and 6th sentences of the 6 sentences in the first text, the labels corresponding to the 3rd and 6th sentences in the first text are recorded as 1, and the labels of the remaining sentences are recorded as 0. The corresponding pseudo labels are [0, 0, 1, 0, 0, 1], indicating that the first text includes 6 sentences, the sentences corresponding to the label 1 are the sentences where the entities of the entity pair are located, and the sentences corresponding to the label 0 do not include the entities of the entity pair.

[0072] Then, the feature representation of each entity of the entity pair is fused to obtain a first feature representation. Before fusion, after obtaining the entity pair from the output sequence, the feature representation of each entity of the entity pair can be determined based on the feature representation of each entity of the entity pair and each word in the above-mentioned first text. That is, at this time, the encoder no longer needs to re-extract features for each entity of the entity pair and the feature representation corresponding to each entity. Instead, it directly obtains the feature representation of each entity of the entity pair based on the feature extraction of the first text and the feature representation of each word in the first text. This is more efficient. Then, the feature representation corresponding to each entity of the entity pair can be spliced ​​(contact) to obtain the first feature representation. Of course, there can be other fusion methods such as weighted average fusion, etc., which are not specifically limited in this application.

[0073] Then, based on the first feature representation and the feature representation of each sentence, the first probability of each sentence being used to prove the relationship between the two entities of the entity pair is predicted, where the first probability corresponding to the nth sentence is shown in formula (6):

[0074]

[0075] Among them, P(Sn |b g ,b h ) represents the first probability corresponding to the nth sentence, Sn represents the nth sentence, b g represents one entity in an entity pair, b h represents the other entity in the entity pair, Represents the first feature representation (it should be noted that Formula 5 is explained by splicing the two entities of the entity pair as an example, that is, ), β, C, and W represent model training parameters.

[0076] Then, based on the first probability corresponding to each sentence and the pseudo-label of the entity pair, the second loss is determined. For example, the label corresponding to each sentence can be determined based on the pseudo-label of the entity pair. For example, if the sentence belongs to the sentence with the label 1 in the pseudo-label of the entity pair, the label corresponding to the sentence is 1, otherwise it is 0. Optionally, the two entities of the entity pair are different, for example, it can be indicated that the entity references of the two entities in the first text are different, or it can be indicated that the positions in the first text are different. Then, based on the first probability corresponding to each sentence and the labels corresponding to the entity pairs with different entities, the second loss can be determined. The second loss is shown in formula (7):

[0077]

[0078] Among them, Loss2 represents the second loss, S n represents the nth sentence, J n Indicates the label corresponding to the nth sentence, P(S n |b g ,b h ) represents the first probability corresponding to the nth sentence; that is, since the number of entity pairs included in the output sequence can be multiple, the multiple entity pairs can be screened based on the positions of the two entities of each entity pair in the first text to obtain one or more second entity pairs of entity pairs with different entities, and then the first probability of each sentence in the A sentences in the first text being used to prove the relationship between the two entities of each second entity pair can be used to determine the second loss in combination with formula (6), which can avoid the model from generating invalid entity pairs and improve the quality of model training.

[0079] Finally, based on the first loss and the second loss, the loss is determined. For example, the first loss and the second loss may be summed to obtain the loss, or the first loss and the second loss may be weighted averaged to obtain the loss, or the first loss and the second loss may be weighted summed to obtain the loss. For example, the loss is as shown in formula (8):

[0080] Loss=Loss1(ω)+σLoss2(8)

[0081] Among them, Loss represents loss, Loss1(ω) represents the first loss, Loss2 represents the second loss, σ represents a hyperparameter, and this application does not specifically limit the value of the hyperparameter.

[0082] S105: Train the first model based on the loss.

[0083] In an embodiment of the present application, after obtaining the loss, iteration or back-propagation training is performed based on the loss until the obtained loss is less than a preset loss threshold, thereby obtaining a trained first model. Then, based on the trained first model, the entity relationships in the input text (whether at the sentence level or the document level) can be extracted, that is, the input text is input into the trained first model to obtain one or more entity pairs, and one or more entity relationships between the two entities of each entity pair.

[0084] For ease of understanding, the following diagrams are used to explain the model training. Figure 2 , Figure 2 A training diagram of a model provided in an embodiment of the present application.

[0085] like Figure 2 As shown, taking the first model as an example, the first model of the present application includes a relationship extraction module and a sentence prediction module. The relationship extraction module includes an encoder and a decoder. The encoder and decoder are not described here. After obtaining the first text, that is, the input sequence, assuming that the first text includes sentence A, the first text is input into the encoder for feature extraction to obtain the feature representation of each word in the first text (i.e. Figure 2 Word Embedding shown), and the feature representation of each sentence in the A sentences in the first text can be obtained based on the feature representation of each word (i.e. Figure 2Sentence Embedding shown); in addition, when the encoder gradually processes each word token in the first text, it updates the initial hidden state of each word in the first text, and finally the encoder outputs the hidden state of each word (or each time step) in the first text, wherein the hidden state of each word in the first text is based on the feature representation of each word and the hidden state of the previous word of each word (for example, obtained by adding or splicing the two, which is not limited in this application), that is, the hidden state of each time step is obtained by forward calculation of a recursive network (such as a recursive neural network RNN, a long short-term memory network LSTM, a gated recurrent unit network GRU, etc., which is not limited in this application) at each time step, which integrates the feature representation of the current word and the context information before the current word, that is, the state of each time step not only represents the information of the current word, but also reflects the context information of the entire input sequence.

[0086] Then, when the decoder starts decoding, it can use the cross attention mechanism to let the decoder focus on different parts of the input sequence at each time step. By calculating the importance of each word token in the input sequence, it generates a context vector and then obtains the probability distribution corresponding to each word position in the output sequence. Specifically:

[0087] Taking the tth time step as an example, at the tth time step (or the tth word unit position in the corresponding output sequence), the decoder first determines the hidden state of the decoder at the tth time step based on the input of the tth time step (including the output of the previous time step, i.e., the t-1th time step, which can also be called the word unit generated in the t-1th time step, and the label of the t-1th time step) and the hidden state of the decoder at the t-1th time step. For example, it is obtained by forward calculation through a recursive network such as a recursive neural network (RNN), a long short-term memory network (LSTM), a recurrent neural network (GRU), etc., which is not limited in this application. When t=1, the input of the first time step includes the feature representation of each word unit in the input sequence or the hidden state of each word unit in the input sequence output by the encoder, and then the hidden state of each subsequent time step includes the feature representation of each word unit in the input sequence.

[0088] Then, based on the hidden state of the decoder at the tth time step and the hidden state of each word unit output by the encoder, the attention score of the hidden state of the decoder at the tth time step and each word unit output by the encoder is determined. The attention score represents the correlation between the hidden state of the decoder at the tth time step and the hidden state of each time step output by the encoder. For example, the attention score of the hidden state of the tth time step and each word unit output by the encoder can be obtained by performing a dot product between the hidden state of the decoder at the tth time step and the hidden state of each word unit output by the encoder, or performing a scaled dot product, or calculating the result by other similarity functions, which is not limited in this application.

[0089] Then, the attention scores of the hidden state of each word unit output by the decoder and the encoder at the tth time step are normalized, for example, by processing with the softmax function, to obtain the attention weight of the hidden state of each word unit output by the encoder for the decoder at the tth time step, that is, the contribution of the hidden state of each word unit output by the encoder to the output of the decoder at the tth time step; then, based on the hidden state of each word unit output by the encoder for the attention weight of the decoder at the tth time step, the hidden state of each word unit output by the encoder is weighted and summed to obtain the context vector corresponding to the decoder at the tth time step.

[0090] The context vector corresponding to the decoder at the tth time step and the hidden state of the decoder at the tth time step are then fused (such as splicing or addition, which is not limited in this application) to obtain the feature vector of the decoder at the tth time step; then based on the feature vector of the decoder at the tth time step, the probability distribution corresponding to the tth time step or the tth word position in the output sequence is obtained, for example, the feature vector of the decoder at the tth time step is mapped to the dimension of the vocabulary through a fully connected layer (Linear Layer) to obtain the score corresponding to the tth time step or the tth word position in the output sequence, and then the score corresponding to the tth time step or the tth word position in the output sequence is normalized, for example, through a softmax function, to obtain the probability distribution corresponding to the score corresponding to the tth time step or the tth word position in the output sequence, and then the probability distribution corresponding to each word position in the output sequence can be obtained by the same token.

[0091] After obtaining the probability distribution corresponding to each word position in the output sequence, the output sequence can be obtained based on the probability distribution corresponding to each word position in the output sequence. The principle will not be repeated here, such as Figure 2 The "Z1Z2Z3 Z4 <t> <r1> <end> ", where the diamond shape corresponds to Z1Z2Z3Z4 which is a word copied from the input sequence. I will not go into details here. The triangle shape corresponds to< / end> < / r1> < / t> <t> <r1> are special tags in the above vocabulary. The principle will not be repeated here. Based on these special tags, we can know the entities and entity relationships in the output sequence. For example, similar to entity 1< / r1> < / t> Entity 2 <t> <r1>, indicating that the entities in the entity pair are entity 1 and entity 2, and their entity relationship is R1. Then, based on the probability distribution of each word position in the output sequence, the entity pair is determined. The principle will not be repeated here.

[0092] After determining the entity pair, the feature representation of each entity in the entity pair can be obtained based on the feature representation of the entity in the entity pair and each word in the first text, and the feature representation of each entity pair in each entity pair can be fused to obtain the first feature representation (i.e. Figure 2 Context Embedding shown in Figure 2); then the first feature representation and the feature representation of each sentence in A sentences are input into the sentence prediction module, and finally the loss is determined based on the output of the sentence prediction module and the output of the decoder in the relation extraction module (i.e. Figure 2 Loss shown), for example: based on the probability distribution of each word position obtained by decoding the decoder in the above-mentioned relationship extraction module and the label corresponding to each word position, a first loss is determined (the principle is not repeated here); and based on the sentence prediction module, a first probability of each sentence in A sentences being used to prove the relationship between the two entities of the entity pair is obtained (the principle is not repeated here), and then based on the first probability corresponding to each sentence in A sentences and the pseudo label of the entity pair, a second loss is determined (the principle is not repeated here); then based on the first loss and the second loss, a loss Loss is determined (the principle is not repeated here), and then the first model is trained based on the loss, and then the entity relationship can be extracted based on the trained first model.

[0093] After training the first model based on the model training method in the above embodiment to obtain the trained first model, the entity relationship extraction device can obtain the second text. Then, the entity relationship extraction device can use the trained first model to extract entity relationships on the second text to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair. The following will specifically describe how the entity relationship extraction device uses the trained first model to extract entity relationships on the second text to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair, as follows:

[0094] See Figure 3 , Figure 3 A flowchart of an entity relationship extraction method provided in an embodiment of the present application. The method includes but is not limited to steps S301-S305:

[0095] S301 : Perform feature extraction on the second text to obtain a feature representation of each word in the second text and a feature representation of each sentence in the second text.

[0096] It should be noted that the principle of step S301 refers to the explanation of the above step S101 and will not be repeated here.

[0097] S302 : Extract entity relationships based on the feature representation of each word in the second text to obtain a second entity pair corresponding to the second text and a first relationship between two entities of the second entity pair.

[0098] In an embodiment of the present application, entity relationship extraction is performed based on the feature representation of each word in the second text, and a probability distribution corresponding to each word position in the output sequence corresponding to the second text is obtained (which will not be repeated here). Then, based on the probability distribution corresponding to each word position in the output sequence, the second entity pair in the second text and the first relationship between the two entities of the second entity pair are determined (which may also be referred to as the first entity relationship extraction result, that is, the first entity relationship extraction result includes the second entity pair and the first relationship corresponding to the second entity pair). The principle will not be repeated here, and you can refer to the corresponding explanation of steps S102-103.

[0099] S303: Determine a first sentence corresponding to the second entity pair based on the feature representation of each sentence in the second text and the feature representation of each entity of the second entity pair.

[0100] In an embodiment of the present application, the first sentence is a sentence in the second text used to prove the relationship between the two entities of the second entity pair. For example, the feature representation of each entity of the second entity pair is fused (such as splicing, adding, etc., which is not limited in this application) to obtain a second feature representation; then, based on the second feature representation and the feature representation of each sentence in the second text, the second probability of each sentence in the second text being used to prove the relationship between the two entities of the second entity pair is predicted, wherein the principle for determining the second probability is similar to the principle for determining the first probability in the above embodiment, and will not be repeated here; then, based on the second probability of each sentence in the second text being used to prove the relationship between the two entities of the second entity pair, the first sentence corresponding to the second entity pair is determined, such as determining the sentence with the largest second probability as the first sentence corresponding to the second entity pair, or determining the sentence with a second probability greater than a preset threshold as the first sentence corresponding to the second entity pair, that is, the number of first sentences corresponding to the second entity pair is one or more, which is not specifically limited in this application.

[0101] S304 : Extract entity relationships based on the feature representation of the first sentence to obtain a third entity pair corresponding to the second text and a second relationship between two entities of the third entity pair.

[0102] In an embodiment of the present application, the second entity relationship extraction result includes a third entity pair in the second text, and a relationship between two entities of the third entity pair. After obtaining the first sentence corresponding to the second entity pair, if the number of the first sentences is one, feature extraction can be performed on the first sentence to obtain a feature representation of the first sentence, and then entity relationship extraction is performed based on the feature representation of the first sentence according to the principle of step S302 to obtain a third entity pair in the second text and a second relationship between the two entities of the third entity pair (similarly, it can also be called a second entity relationship extraction result. In this case, the second entity relationship extraction result includes the third entity pair and the second relationship between the two entities of the third entity pair); similarly, if the number of the first sentences is multiple, the multiple first sentences can be combined to obtain a combined text. For example, the multiple first sentences can be combined and spliced ​​in order based on their positions in the second text to obtain a combined text, or the multiple first sentences can be randomly combined and spliced ​​to obtain a combined text. This application does not make specific restrictions. Feature extraction is then performed on the combined text to obtain a feature representation of each word in the combined text. Entity relationship extraction is then performed based on the feature representation of each word in the combined text according to the principle of step S302 to obtain a third entity pair in the second text and a second relationship between the two entities of the third entity pair, that is, the second entity relationship extraction result.

[0103] S305 : Fusing the second entity pair, the first relationship, the third entity pair, and the second relationship to obtain the first entity pair and the relationship between the two entities of the first entity pair.

[0104] In an embodiment of the present application, the number of the second entity pairs and the third entity pairs can be one or more, and the first relationship between the two entities of each second entity pair and the second relationship between the two entities of each third entity pair can also be one or more. This application does not make specific limitations. That is to say, after obtaining the first entity relationship extraction result and the second entity relationship extraction result, that is, after obtaining the second entity pair, the first relationship, the third entity pair, and the second relationship, the second entity pair, the first relationship, the third entity pair, and the second relationship can be fused, that is, the first entity relationship extraction result and the second entity relationship extraction result can be fused to obtain the first entity pair and the relationship between the two entities of the first entity pair (that is, to obtain the final entity relationship extraction result, which includes the first entity pair and the relationship between the two entities of the first entity pair). Specifically:

[0105] If the second entity pair and the third entity pair are the same, the second entity pair or the third entity pair is determined as the first entity pair, and the intersection or union of the first relationship and the second relationship is taken to obtain the relationship between the two entities of the first entity pair; if the second entity pair and the third entity pair are different, the second entity pair and the third entity pair are both determined as the first entity pair, and the first relationship is determined as the relationship between the two entities of the second entity pair included in the first entity pair, and the second relationship is determined as the relationship between the two entities of the third entity pair included in the first entity pair, which can improve the accuracy of entity relationship extraction.

[0106] It should be noted that this embodiment is described by taking the number of the second entity pair and the third entity pair as one as an example. The second entity pair and the third entity pair being the same means that the two entities of the second entity pair are exactly the same as the two entities of the third entity pair, that is, for example, if the two entities of the second entity pair are entity 1 and entity 2, then the two entities of the third entity pair must also be entity 1 and entity 2 in order to consider that the second entity pair and the third entity pair are the same; when the number of the second entity pair or the third entity pair is multiple, the second entity pair and the third entity pair being the same means that the fourth entity pair (the number is one or more) of the multiple second entity pairs is the same as the third entity pair, or that the fifth entity pair (the number is one or more) of the multiple third entity pairs is the same as the second entity pair; when the number of both the second entity pair and the third entity pair is multiple, the second entity pair and the third entity pair being the same means that the fourth entity pair of the multiple second entity pairs is the same as the fifth entity pair of the multiple third entity pairs, and then the same processing is performed according to the same principle as the second entity pair and the third entity pair. The principles are similar and will not be repeated here. On the contrary, the interpretation of the difference between the second entity pair and the third entity pair is similar to the interpretation of the sameness between the second entity pair and the third entity pair, and will not be repeated here.

[0107] For example, assume that the second entity pair and the first relationship are as follows: Entity 1 < / t> Entity 2 <t> R1, Entity 3< / t> Entity 4 <t> R2, Entity 4< / t> Entity 5 <t> R4, which indicates that the second entity pair includes the second entity pair (1) consisting of entity 1 and entity 2, the second entity pair (2) consisting of entity 3 and entity 4, and the second entity pair (3) consisting of entity 4 and entity 5, where R1 is the first relationship between entity 1 and entity 2, R2 is the first relationship between entity 3 and entity 4, and R4 is the first relationship between entity 4 and entity 5. Similarly, assuming that the examples of the third entity pair and the second relationship are as follows: entity 1< / t> Entity 2 <t> R1R3, Entity 3< / t> Entity 4 <t>R2, which indicates that the third entity pair includes the third entity pair (1) consisting of entity 1 and entity 2, and the third entity pair (2) consisting of entity 3 and entity 4, wherein R1 and R3 are both the second relationship between entity 1 and entity 2, and R2 is the second relationship between entity 3 and entity 4; at this time, the number of the second entity pair and the third entity pair are both multiple, and the second entity pair (1) in the second entity pair is the same as the third entity pair (1) in the third entity pair, then the second entity pair (1) or the third entity pair (1) is determined as the first entity pair, and the first relationship of the second entity pair (1), i.e. R1, and the second relationship of the third entity pair (1), i.e. R1R3, are taken as the union or intersection, and the relationship between the two entities of the first entity pair is obtained as R 1R3 or R1; and the second entity pair (2) in the second entity pair is the same as the third entity pair (2), then the second entity pair (2) or the third entity pair (2) is determined to be the first entity pair, and the first relationship of the second entity pair (2), i.e., R2, and the second relationship of the third entity pair (2), i.e., R2, are taken as a union or an intersection, and the relationship between the two entities of the first entity pair is respectively obtained to be R2 or a null value (i.e., the intersection is null in this case); and the second entity pair (3) in the second entity pair is different from the third entity pair (1) and the third entity pair (2) in the third entity pair, then the second entity pair (3) is determined to be the first entity pair, and the first relationship R4 of the second entity pair (3) is determined to be the relationship between the two entities of the first entity pair.

[0108] It should be noted that the entity relationship extraction device uses the trained first model in addition to executing Figure 3 In addition to steps S301-S305 in the entity relationship extraction method shown, other steps in the above embodiment can also be performed, and the same technical effect can be achieved; of course, Figure 3 The entity relationship extraction method shown can also be executed by the first model obtained by training in the above embodiment, which will not be repeated here.

[0109] To further facilitate understanding, the following example illustrates entity relationship extraction using the trained first model. Figure 4 , Figure 4 This is a schematic diagram of entity relationship extraction provided in an embodiment of the present application. Figure 4 As shown, the first model includes a relation extraction module and a sentence prediction module. The relation extraction module includes an encoder and a decoder. Assume that the second text includes M sentences, then the second text is input into the encoder, and the feature representation of each word in the second text is output, that is, Figure 4 The Word Embedding (1) shown in FIG5 is then input into the decoder for entity relationship extraction to obtain the second entity pair and the first relationship between the two entities of the second entity pair, that is, Figure 4 The first entity relationship extraction result shown includes entity 1< / t> Entity 2 <t> R1, Entity 3< / t> Entity 4 <t>R2, which indicates that the second entity pair in the first entity relationship extraction result includes the second entity pair (1) composed of entity 1 and entity 2, and the second entity pair (2) composed of entity 3 and entity 4, where R1 is the first relationship between entity 1 and entity 2, and R2 is the first relationship between entity 3 and entity 4; then, based on the feature representation of each word in the second entity pair and the second text, the feature representation of each entity in the second entity pair is determined, and the feature representation of each entity in the second entity pair is fused to obtain the second feature representation (i.e. Figure 4 Context Embedding shown), it should be noted that, if the number of second entity pairs is multiple, then the feature representation of each entity of each second entity pair is determined in the same way, and then the feature representation of each entity of each second entity pair is fused to obtain the fused feature corresponding to each second entity pair, and then the fused features corresponding to multiple second entity pairs are fused (such as splicing, addition, etc., which are not limited in this application) to obtain the second feature representation; and based on the feature representation of each word in the second text, the feature representation of each sentence in the M sentences in the second text is determined (i.e. Figure 4 The second feature representation and the feature representation of each sentence in the M sentences in the second text are then input into the sentence prediction module, and the first sentence corresponding to the second entity pair is output; the first sentence (or the combined text when there are multiple first sentences) is then input into the encoder to obtain the feature representation of each word in the first sentence, that is, Figure 4 The Word Embedding (1) shown in the figure is then input into the decoder to extract the entity relationship, and the third entity pair and the second relationship between the two entities of the third entity pair are obtained. Figure 4 The second entity relationship extraction result shown includes entity 1< / t> Entity 2 <t> R1R3, Entity 3< / t> Entity 4 <t>R2, which indicates that the third entity pair in the second entity relationship extraction result includes the third entity pair (1) composed of entity 1 and entity 2, and the third entity pair (2) composed of entity 3 and entity 4, where R1 and R3 are both the second relationship between entity 1 and entity 2, and R2 is the second relationship between entity 3 and entity 4; then the second entity pair, the first relationship, the third entity pair, and the second relationship are fused, that is, the first entity relationship extraction result and the second entity relationship extraction result are fused to obtain the first entity pair and the relationship between the two entities of the first entity pair, that is, Figure 4 The entity relationship extraction results shown in FIG. 1 include a first entity pair (1) consisting of entity 1 and entity 2, and a first entity pair (2) consisting of entity 3 and entity 4. The relationship between the two entities in the first entity pair (1) includes R1 and R3, and the relationship between the two entities in the first entity pair (2) includes R2. It should be explained that Figure 4 The entity relationship extraction shown is only an example in the above embodiment and is not limited. We will not elaborate here on how to fuse the first entity pair and the relationship between the two entities of the first entity pair. It should be noted that the specific implementation principles of each step in this embodiment have been explained above and will not be repeated here.

[0110] Of course, in an optional embodiment, the first model may also include multiple identical relationship extraction modules, and the encoders and decoders in any two relationship extraction modules are the same. Then, multiple first texts can be input into multiple relationship extraction modules respectively to obtain entity relationship extraction results corresponding to each relationship extraction module, wherein the input data of each relationship extraction module, i.e., the first text, can be the same or different, and then the entity relationship extraction results corresponding to each relationship extraction module are fused (such as average or weighted average, etc.) to obtain the final relationship extraction result; or, the first model may also include multiple different relationship extraction modules, and the encoders and decoders in any two relationship extraction modules are different, then the first texts can be input into the multiple relationship extraction modules respectively to obtain entity relationship extraction results corresponding to each relationship extraction module, and then the entity relationship extraction results corresponding to each relationship extraction module are fused (such as average or weighted average, etc.) to obtain the final relationship extraction result. That is to say, the use of integrated learning can reduce the deviation of a single model, improve stability and generalization ability, and further improve model performance and robustness to enhance the predictive ability of the model.

[0111] See Figure 5 , Figure 5 A schematic diagram of an entity relationship extraction system provided in an embodiment of the present application.

[0112] Figure 5 The system shown includes an entity relationship extraction device and a client; the entity relationship extraction device can be a server, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as basic cloud computing services such as big data and artificial intelligence platforms, which are not specifically limited in this application; the client can be a smart phone, tablet computer, laptop computer, desktop computer, smart TV, desktop computer, smart watch, smart car and other smart terminals, but is not limited to these.

[0113] The user can upload the second text through the client, and then the client sends the second text to the entity relationship extraction device, and then the entity relationship extraction device performs entity relationship extraction on the second text, that is, performs entity relationship extraction on the second text through the trained first model to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair, specifically as follows: perform feature extraction on the second text to obtain the feature representation of each word in the second text and the feature representation of each sentence in the second text; then perform entity relationship extraction based on the feature representation of each word in the second text to obtain the first entity relationship extraction result corresponding to the second text, wherein the first entity relationship extraction result includes the second entity pair corresponding to the second text and the first relationship between the two entities of the second entity pair; then determine the entity relationship with the second entity pair based on the feature representation of each sentence in the second text and the feature representation of each entity of the second entity pair. The first sentence corresponding to the entity pair, wherein the first sentence is a sentence in the second text used to prove the relationship between the two entities of the second entity pair; then, entity relationship extraction is performed based on the feature representation of the first sentence to obtain a second entity relationship extraction result corresponding to the second text, wherein the second entity relationship extraction result includes a third entity pair corresponding to the second text and a second relationship between the two entities of the third entity pair; finally, based on the first entity relationship extraction result and the second entity relationship extraction result, the entity relationship extraction result corresponding to the second text is determined, or in other words, the second entity pair, the first relationship, the third entity pair and the second relationship are integrated to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair; further, the entity relationship extraction device sends the entity relationship extraction result to the client, that is, sends the first entity pair of the second text and the relationship between the two entities of the first entity pair to the client.

[0114] It should be noted that Figure 5 The entity relationship extraction device in the embodiment can also correspondingly execute the steps executed by the entity relationship extraction device in the above embodiment, which will not be repeated here.

[0115] See Figure 6 , Figure 6 A schematic diagram of a model training system provided in an embodiment of the present application.

[0116] Figure 6 The system shown includes a model training device and a client; the model training device can be a server, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as basic cloud computing services such as big data and artificial intelligence platforms, which are not specifically limited in this application; the client can be a smart phone, tablet computer, laptop computer, desktop computer, smart TV, desktop computer, smart watch, smart car and other smart terminals, but is not limited to these.

[0117] The user can upload the first text through the client, and then the client sends the first text to the model training device, and then the model training device performs model training based on the first text, specifically as follows: the model training device extracts features from the first text to obtain feature representations of each word in the first text and feature representations of each sentence in the first text; then the model training device extracts entity relationships based on the feature representations of each word to obtain the probability distribution of each word position in the output sequence; based on the probability distribution of each word position, the entity pairs in the first text are determined; then the model training device determines the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence; then the model training device trains the first model based on the loss.

[0118] It should be noted that Figure 6 The model training device in the embodiment can also perform the steps performed by the model training device in the above embodiment, which will not be repeated here.

[0119] It should be noted that the main scenarios applicable to the entity relationship extraction of this application may include but are not limited to the following scenarios:

[0120] (1) Information extraction: extracting key information from unstructured text, such as the relationships between people, places, and events in news reports; (2) Knowledge graph construction: creating a structured knowledge base to enhance the intelligence of search engines and recommendation systems; (3) Question-answering system: understanding the entities and relationships in user questions and improving the accuracy of the system's answers; (4) Social network analysis: analyzing the relationships between users, performing social network mining and sentiment analysis; (5) Medical data analysis: extracting the relationships between diseases, drugs, and symptoms from medical literature to assist medical research and diagnosis, etc.

[0121] It can be seen that in an embodiment of the present application, by training a model, specifically: feature extraction is performed on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text; and then entity relationship extraction is performed based on the feature representation of each word, so that the probability distribution of each word position in the output sequence can be obtained, that is, based on the feature representation of each word, each word position in the output sequence obtained by performing entity relationship extraction includes entity pairs and word positions corresponding to the relationship between entities in the entity pairs, then the entity pairs in the first text can be determined based on the probability distribution of each word position; and then the loss is determined based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence in the first text; and then the first model is trained based on the loss, that is, since each word position in the output sequence obtained includes entity pairs and word positions corresponding to the relationship between entities in the entity pairs, that is, the probability distribution of word positions corresponding to the entity pairs, the entity pairs, The probability distribution of the word position corresponding to the relationship between the entities in the pair. After obtaining the probability distribution of each word position, the loss is comprehensively calculated by combining the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence in the first text. Compared with directly determining the entity pair and the corresponding entity relationship based on the probability distribution of each word position and then determining the loss, the feature representation of each entity in the entity pair and the feature representation of each sentence in the first text are added to determine the loss. By combining the feature representation of each entity in the entity pair and the feature representation of each sentence, the model can learn multi-dimensional considerations and enhance the evidence for the extraction of entity relationships, so that when the trained first model is used for entity relationship extraction, the accuracy of entity relationship extraction can be improved, and the first probability of each sentence belonging to the first sentence is predicted directly based on the feature representation of the entity and the feature representation of each sentence, avoiding the interference of noise on model training, which can improve the quality of model training and thus improve the accuracy of entity relationship extraction.

[0122] In addition, the present application also extracts features from the second text to obtain feature representations of each word unit in the second text and feature representations of each sentence in the second text; then, entity relationship extraction is performed based on the feature representation of each word unit to obtain a second entity pair in the second text and the relationship between the two entities of the second entity pair; then, based on the feature representation of each sentence in the second text and the feature representation of each entity of the second entity pair, the first sentence corresponding to the second entity pair is determined, wherein the first sentence is a sentence in the second text used to prove the relationship between the two entities of the second entity pair; then, entity relationship extraction is performed based on the feature representation of the first sentence to obtain a third entity pair in the second text and the relationship between the two entities of the third entity pair; finally, based on the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, and the relationship between the two entities of the third entity pair The relationship between entities is determined to determine the entity relationship extraction result corresponding to the second text. That is to say, after performing entity relationship extraction based on the feature representation of each word unit to obtain the second entity pair in the second text and the relationship between the two entities of the second entity pair, the feature representation of each sentence and the feature representation of each entity of the second entity pair are combined to determine the first sentence corresponding to the second entity pair. This can avoid the introduction of noise and improve the accuracy of determining the first sentence. Then, based on the first sentence, the third entity pair and the relationship between the two entities of the third entity pair are predicted to improve the accuracy of entity relationship extraction. Finally, the second entity pair, the relationship between the two entities of the second entity pair, the third entity pair, and the relationship between the two entities of the third entity pair are combined to determine the final entity relationship extraction result. Compared with single-dimensional entity relationship extraction, it has higher accuracy.

[0123] See Figure 7 , Figure 7 This is a block diagram of the functional units of a model training device provided in an embodiment of the present application. The model training device 700 includes: a first acquisition unit 701 and a first processing unit 702;

[0124] A first acquiring unit 701 is configured to acquire a first text;

[0125] A first processing unit 702 is configured to perform feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text;

[0126] The first processing unit 702 is further configured to extract entity relationships based on the feature representation of each word, and obtain a probability distribution of each word position in the output sequence;

[0127] The first processing unit 702 is further configured to determine entity pairs in the first text based on the probability distribution of each word position;

[0128] The first processing unit 702 is further configured to determine a loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence;

[0129] The first processing unit 702 is further configured to train the first model based on the loss.

[0130] In one embodiment of the present application, in determining the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence, the first processing unit 702 is specifically configured to:

[0131] Determine a first loss based on the probability distribution of each word position and the label corresponding to each word position;

[0132] determining a second loss based on a feature representation of each entity of the entity pair, a feature representation of each sentence, and a pseudo-label of the entity pair, wherein the pseudo-label is used to indicate a sentence in the first text that proves the relationship between the two entities of the entity pair;

[0133] The loss is determined based on the first loss and the second loss.

[0134] In one embodiment of the present application, in determining the second loss based on the feature representation of each entity in the entity pair, the feature representation of each sentence, and the pseudo-label of the entity pair, the first processing unit 702 is specifically configured to:

[0135] Fusing the feature representations of each entity in the entity pair to obtain a first feature representation;

[0136] Based on the first feature representation and the feature representation of each sentence, predicting a first probability for each sentence to prove the relationship between two entities of the entity pair;

[0137] The second loss is determined based on the first probability and pseudo label of the entity pair corresponding to each sentence.

[0138] In one embodiment of the present application, the first processing unit 702 is specifically configured to:

[0139] Based on the sentence where each entity of the entity pair in the first text is located, a pseudo label of the entity pair is determined.

[0140] In one embodiment of the present application, the two entities of the entity pair are located at different positions in the first text.

[0141] In one embodiment of the present application, in determining the first loss based on the probability distribution of each word position and the label corresponding to each word position, the first processing unit 702 is specifically configured to:

[0142] Based on the probability distribution of each word position, determine the conditional probability of each word position under the corresponding label;

[0143] The first loss is then determined based on the conditional probability of each word position under the corresponding label.

[0144] In a specific implementation, the first acquisition unit 701 and the first processing unit 702 described in the embodiment of the present invention can also execute other implementation methods described in the embodiment of the training method of the model provided by the embodiment of the present invention, which will not be repeated here.

[0145] See Figure 8 , Figure 8 This is a block diagram of the functional units of an entity relationship extraction device provided in an embodiment of the present application. The entity relationship extraction device 800 includes: a second acquisition unit 801 and a second processing unit 802;

[0146] A second acquiring unit 801 is configured to acquire a second text;

[0147] The second processing unit 802 is configured to extract entity relationships from the second text using the first model trained in the above embodiment to obtain a first entity pair of the second text and a relationship between two entities of the first entity pair.

[0148] In one embodiment of the present application, in extracting entity relationships from the second text to obtain a first entity pair of the second text and a relationship between two entities of the first entity pair, the second processing unit 802 is specifically configured to:

[0149] Performing feature extraction on the second text to obtain a feature representation of each word in the second text and a feature representation of each sentence in the second text;

[0150] Performing entity relationship extraction based on the feature representation of each word in the second text to obtain a second entity pair corresponding to the second text and a first relationship between two entities of the second entity pair;

[0151] Determining a first sentence corresponding to the second entity pair based on a feature representation of each sentence in the second text and a feature representation of each entity in the second entity pair, wherein the first sentence is a sentence in the second text used to prove the relationship between two entities of the second entity pair;

[0152] Performing entity relationship extraction based on the feature representation of the first sentence to obtain a third entity pair corresponding to the second text and a second relationship between two entities of the third entity pair;

[0153] The second entity pair, the first relationship, the third entity pair, and the second relationship are fused to obtain the first entity pair and the relationship between the two entities of the first entity pair.

[0154] In one embodiment of the present application, in terms of fusing the second entity pair, the first relationship, the third entity pair, and the second relationship to obtain the first entity pair and the relationship between the two entities of the first entity pair, the second processing unit 802 is specifically configured to:

[0155] If the second entity pair and the third entity pair are the same, determining the second entity pair or the third entity pair as the first entity pair, and taking the intersection or union of the first relationship and the second relationship to obtain the relationship between the two entities of the first entity pair;

[0156] When the second entity pair and the third entity pair are different, the second entity pair and the third entity pair are both determined as the first entity pair, and the first relationship is determined as the relationship between the two entities of the second entity pair in the first entity pair, and the second relationship is determined as the relationship between the two entities of the third entity pair in the first entity pair.

[0157] In a specific implementation, the second acquisition unit 801 and the second processing unit 802 described in the embodiment of the present invention may also execute other implementation methods described in the embodiment of the entity relationship extraction method provided by the embodiment of the present invention, which will not be repeated here.

[0158] See Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 9 As shown, electronic device 900 includes a transceiver 901, a processor 902, and a memory 903. These are connected via a bus 904. The memory 903 is used to store computer programs and data, and can transmit the data stored in the memory 903 to the processor 902.

[0159] The electronic device 900 may be a model training device 700 or an entity relationship extraction device 800;

[0160] When the electronic device 900 is a model training device 700, the processor 902 is configured to read the computer program in the memory 903 and perform the following operations:

[0161] Controlling the transceiver 901 to obtain a first text;

[0162] Performing feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text;

[0163] Entity relations are extracted based on the feature representation of each word, and the probability distribution of each word position in the output sequence is obtained;

[0164] Determining entity pairs in the first text based on the probability distribution of each word position;

[0165] Determine the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence;

[0166] The first model is trained based on the loss.

[0167] In a specific implementation, the transceiver 901 and processor 902 described in the embodiment of the present invention may also execute other implementation methods described in the embodiment of the training method of the model provided by the embodiment of the present invention, which will not be repeated here.

[0168] When the electronic device 900 is the entity relationship extraction device 800, the processor 902 is configured to read the computer program in the memory 903 and perform the following operations:

[0169] Controlling the transceiver 901 to obtain a second text;

[0170] The first model trained in the above embodiment is used to extract entity relationships from the second text to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair.

[0171] In a specific implementation, the transceiver 901 and the processor 902 described in the embodiment of the present invention may also execute other implementation methods described in the embodiment of the entity relationship extraction provided by the embodiment of the present invention, which will not be repeated here.

[0172] Specifically, the transceiver 901 may be Figure 7 The first acquisition unit 701 of the training device 700 of the model of the embodiment or Figure 8 The second acquisition unit 801 of the entity relationship extraction device 800 of the embodiment, the processor 902 can be Figure 7 The first processing unit 702 of the training device 700 of the model of the embodiment or Figure 8 The second processing unit 802 of the entity relationship extraction device 800 of the embodiment.

[0173] It should be understood that the electronic device in this application can be a model training device or an entity relationship extraction device, and both the model training device and the entity relationship extraction device can be terminal devices or servers, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms. This application does not make specific restrictions. The above-mentioned electronic devices are only examples, not exhaustive, and include but are not limited to the above-mentioned electronic devices.

[0174] It should be understood that the embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any model training method or entity relationship extraction method recorded in the above method embodiments.

[0175] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps of any model training method or entity relationship extraction method recorded in the above method embodiments.

[0176] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0177] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0178] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0179] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0180] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of software program modules.

[0181] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0182] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0183] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.< / t>

Claims

1. A model training method, characterized in that: The method comprises: Performing feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text; Entity relations are extracted based on the feature representation of each word, and the probability distribution of each word position in the output sequence is obtained; Determining entity pairs of the first text based on the probability distribution of each word position; Determining a loss based on a probability distribution of each word position, a feature representation of each entity in the entity pair, and a feature representation of each sentence; A first model is trained based on the loss.

2. The method according to claim 1, characterized in that The determining of the loss based on the probability distribution of each word position, the feature representation of each entity in the entity pair, and the feature representation of each sentence includes: Determine a first loss based on the probability distribution of each word position and the label corresponding to each word position; Determining a second loss based on a feature representation of each entity of the entity pair, a feature representation of each sentence, and a pseudo-label of the entity pair, wherein the pseudo-label is used to indicate a labeling of an entity of the entity pair for each sentence in the first text; The loss is determined based on the first loss and the second loss.

3. The method according to claim 2, characterized in that The determining of the second loss based on the feature representation of each entity of the entity pair, the feature representation of each sentence, and the pseudo label of the entity pair includes: Fusing the feature representations of each entity in the entity pair to obtain a first feature representation; Based on the first feature representation and the feature representation of each sentence, predicting a first probability that each sentence is used to prove the relationship between two entities of the entity pair; The second loss is determined based on the first probability corresponding to each sentence and the pseudo label of the entity pair.

4. The method according to claim 2 or 3, characterized in that The method further comprises: Based on the sentence where each entity of the entity pair in the first text is located, a pseudo label of the entity pair is determined.

5. The method according to any one of claims 2 to 4, characterized in that: The determining of the first loss based on the probability distribution of each word unit position and the label corresponding to each word unit position includes: Based on the probability distribution of each word position, determine the conditional probability of each word position under the corresponding label; The first loss is determined based on the conditional probability of each word position under the corresponding label.

6. A method for extracting entity relationships, characterized in that: The method comprises: Get the second text; Using the trained first model, entity relationships are extracted from the second text to obtain a first entity pair of the second text and a relationship between two entities of the first entity pair, wherein the trained first model is obtained by training using the training method described in any one of claims 1-5.

7. The method according to claim 6, characterized in that The extracting entity relationships from the second text to obtain a first entity pair of the second text and a relationship between two entities of the first entity pair includes: performing feature extraction on the second text to obtain a feature representation of each word in the second text and a feature representation of each sentence in the second text; Extracting entity relationships based on the feature representation of each word in the second text to obtain a second entity pair corresponding to the second text and a first relationship between two entities in the second entity pair; Determining, based on a feature representation of each sentence in the second text and a feature representation of each entity in the second entity pair, a first sentence corresponding to the second entity pair, wherein the first sentence is a sentence in the second text used to prove the relationship between two entities of the first entity pair; Performing entity relationship extraction based on the feature representation of the first sentence to obtain a third entity pair corresponding to the second text and a second relationship between two entities of the third entity pair; The second entity pair, the first relationship, the third entity pair, and the second relationship are fused to obtain the first entity pair and the relationship between two entities of the first entity pair.

8. The method according to claim 7, characterized in that The fusing the second entity pair, the first relationship, the third entity pair, and the second relationship to obtain the first entity pair and the relationship between two entities of the first entity pair includes: If the second entity pair and the third entity pair are the same, determining the second entity pair or the third entity pair as the first entity pair, and taking the intersection or union of the first relationship and the second relationship to obtain the relationship between the two entities of the first entity pair; When the second entity pair and the third entity pair are different, the second entity pair and the third entity pair are both determined to be the first entity pair, and the first relationship is determined to be the relationship between the two entities of the second entity pair included in the first entity pair, and the second relationship is determined to be the relationship between the two entities of the third entity pair included in the first entity pair.

9. A model training device, characterized in that: The model training device includes: a first acquisition unit and a first processing unit; The first acquiring unit is configured to acquire a first text; The first processing unit is configured to perform feature extraction on the first text to obtain a feature representation of each word in the first text and a feature representation of each sentence in the first text; The first processing unit is further configured to extract entity relationships based on the feature representation of each word element to obtain a probability distribution of each word element position; The first processing unit is further configured to determine entity pairs of the first text based on the probability distribution of each word position; The first processing unit is further configured to determine a loss based on a probability distribution of each word position, a feature representation of each entity of the entity pair, and a feature representation of each sentence; The first processing unit is further used to train the first model based on the loss.

10. An entity relationship extraction device, characterized in that: The entity relationship extraction device includes: a second acquisition unit and a second processing unit; The second acquiring unit is configured to acquire a second text; The second processing unit is used to extract entity relationships from the second text using the trained first model to obtain the first entity pair of the second text and the relationship between the two entities of the first entity pair, wherein the trained first model is obtained by training using the training method described in any one of claims 1 to 5.

11. An electronic device, characterized in that: include: A processor and a memory, the processor is connected to the memory, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Context vector generation method, electronic equipment, storage medium and program product

    CN121052385A