A Dialogue Relationship Extraction Method Incorporating Location-Aware Refinement

By constructing heterogeneous mention dialogue diagrams and introducing position-aware refinement and auxiliary tasks, the interference caused by similar structural information of dialogue text in the prior art is solved, and the accuracy and performance of relationship extraction are improved.

CN115455197BActive Publication Date: 2025-06-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211066885.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-06-27
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

The interference caused by similar structural information of dialogue text in the prior art cannot be classified. The uniqueness of dialogue data is not considered, resulting in low accuracy in relation extraction.

Method used

The dialogue relationship extraction method with fusion position perception is adopted, and the relationship extraction performance is enhanced by constructing a heterogeneous mention dialogue graph, using the graph attention network to update the node characteristics, and auxiliary tasks such as speaker prediction and trigger word prediction are introduced to enhance the relationship extraction performance.

Benefits of technology

Improve the accuracy of relationship extraction, and enhance the performance of dialogue relationship extraction by capturing the unique characteristics of the speaker and discovering highly supported information of the correct relationship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115455197B_ABST
    Figure CN115455197B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of natural language processing, and discloses a dialogue relation extraction method integrating position-aware refinement, which solves the problem that the interference caused by similar structural information in dialogue texts in the prior art makes it impossible to classify nodes and the relationship extraction accuracy is not high due to the failure to consider the uniqueness of dialogue data. The method of the present invention is as follows: First, based on the syntactic analysis of the dialogue, the mention words of the entities are obtained; then, based on the mention words and dialogue information, a heterogeneous mention dialogue graph is constructed and the node features are initialized; then, by using a position-aware refinement graph attention network on the heterogeneous mention dialogue graph, the updated node features are obtained, and the entity dialogue graph is obtained by merging the nodes; finally, by fusing the path information between entity pairs in the entity dialogue graph, the relationship between entity pairs is inferred.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and particularly to a method for extracting dialogue relationships by fusing position perception refinement. Background Art

[0002] The task of relationship extraction aims to identify named entities in unstructured text sources and mine the semantic relationships between entity pairs, realizing the mapping from unstructured text corpora to structured knowledge. Relationship extraction is a core task in the automatic construction process of large-scale knowledge graphs, and also provides knowledge support for applications such as question answering systems and search engines. In current work, more practically significant is relationship extraction based on dialogue. Dialogue relationship extraction refers to extracting relationships from multi-turn conversations among a group of speakers, where relationships exist not only between entities in the dialogue text but also between speakers in each dialogue.

[0003] Existing research often solves this task by methods based on the pre-trained BERT model and graph-based methods. Methods based on pre-trained models often use special tokens to extract entity features in text to highlight information related to entities. Graph-based methods model speakers, entities, and discourse nodes by constructing a graph, encoding long-distance dependencies between parameters, and considering rich semantic information in the dialogue.

[0004] However, for graph-based methods, when two nodes in the graph structure have similar neighborhood topologies, similar feature representations will be learned by the graph neural network, and the nodes cannot be classified. Moreover, the above methods in the prior art regard dialogue as pure text, without considering the uniqueness of dialogue data, such as the trigger word support information and speaker feature information of relationships in the dialogue, and the accuracy of relationship extraction needs to be improved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to propose a method for extracting dialogue relationships by fusing position perception refinement, so as to solve the interference caused by similar structural information in dialogue text in the prior art, the inability to classify nodes, and the problem of low accuracy of relationship extraction caused by not considering the uniqueness of dialogue data.

[0006] The technical solution adopted by the present invention to solve the above technical problem is:

[0007] A method for extracting dialogue relationships by fusing position perception refinement, comprising the following steps:

[0008] A. Training a dialogue relationship extraction model:

[0009] A1. Inputting the dialogue for training and the entity pairs for which relationships need to be extracted;

[0010] A2. Tokenize the original text of the dialogue and perform dependency syntactic analysis. Then, based on the obtained dependency syntactic analysis results, take the words with the same text as the entities in the entity pair and the words with a referential relationship as mention words; Based on the dialogue, segment the utterance sentences according to the change of the speaker or the pause in the dialogue;

[0011] A3. Construct a heterogeneous mention dialogue graph:

[0012] The nodes of the heterogeneous mention dialogue graph include mention nodes, utterance nodes, and dialogue nodes, and the edges include mention dependency edges, speaker dependency edges, utterance dependency edges, and dialogue dependency edges, which are constructed according to the following definitions:

[0013] Mention nodes: Set each mention word as a mention node, and set mention words with the same text but different positions in the dialogue as different mention nodes; Different positions in the dialogue include that mention words with the same text appear in different utterance sentences and that mention words with the same text appear in different positions in the same utterance sentence;

[0014] Utterance nodes: Set each utterance sentence as an utterance node;

[0015] Dialogue nodes: Set the dialogue as a dialogue node;

[0016] The edges formed by connecting each pair of mention nodes corresponding to the same entity constitute mention dependency edges;

[0017] The edges formed by connecting each pair of utterance nodes corresponding to the same speaker, and the edges formed by connecting each pair of mention nodes corresponding to the same speaker and different entities, constitute speaker dependency edges;

[0018] The edges between a mention node and the utterance node corresponding to the utterance sentence to which the corresponding word of the mention node belongs constitute utterance dependency edges;

[0019] The edges between a mention node and the dialogue node constitute dialogue dependency edges;

[0020] A4. Initialize the node features in the heterogeneous mention dialogue graph;

[0021] A5. Based on the position information of the corresponding words of each mention node in the corresponding utterance sentence, and the feature representations of the neighbor nodes of each mention node except the dialogue node, update the feature representations of each mention node in the heterogeneous mention dialogue graph based on the graph attention network; According to the position information of the utterance sentences corresponding to each utterance node in the dialogue, and the feature representations of the neighbor nodes of each utterance node except the dialogue node, update the feature representations of each utterance node in the heterogeneous mention dialogue graph based on the graph attention network;

[0022] A6. Construct an entity dialogue graph:

[0023] Based on the feature representations of the mentioned nodes and discourse nodes obtained by updating, entity nodes and their feature representations are obtained by merging according to the corresponding relationships between the mentioned nodes and entities, and speaker nodes and their feature representations are obtained by merging according to the corresponding relationships between the discourse nodes and speakers;

[0024] Retain the edges in the heterogeneous mention dialogue graph except those eliminated due to node merging, and construct the feature representation of the edge based on the feature representations of the nodes at both ends of the edge;

[0025] A7. Integrate the path information between entity pairs in the integrated entity dialogue graph, and perform reasoning to obtain the relationships between entity pairs;

[0026] A8. Loop through steps A1 - A7, and train the dialogue relationship extraction model based on the binary cross - entropy loss for relation extraction until the model converges or reaches the set number of training rounds;

[0027] B. Perform the dialogue relationship extraction task:

[0028] Use the trained dialogue relationship extraction model to obtain the relationship between the entities of the entity pair with the dialogue text and the entity pair of the relationship to be extracted as the input.

[0029] Further, in step A2, the original text of the dialogue is tokenized and dependency syntactic analysis is performed, specifically including:

[0030] First, remove the non - text part of the original text and perform format encoding; then, use natural language processing tools to tokenize the text and perform dependency syntactic analysis to obtain the dependency syntactic analysis tree.

[0031] Further, in step A4, initialize the node features in the heterogeneous mention dialogue graph, including:

[0032] A41. For each discourse sentence of the dialogue, first, concatenate the speaker label and the sequence composed of the words of the discourse sentence in sequence to form the discourse sentence word sequence, and then, concatenate the sequence composed of the start marker of the dialogue, each discourse sentence word sequence in sequence, and the end marker of the dialogue to form the dialogue word sequence; input the obtained dialogue word sequence into the pre - trained language model to obtain the initial word embedding representation;

[0033] Then, according to the obtained initial word embedding representation, for each discourse sentence, respectively, concatenate the embedding representation i of its speaker s with the initial embedding representation ij of each of its words x to obtain the embedding representation w ij of each of its words:

[0034]

[0035] Among them, represents the splicing operation; w ij represents the embedding representation of the j-th word in the utterance sentence u i ;

[0036] A42. Based on the embedding representations w ij of each word obtained in step A41, the initial features of each mention node are obtained based on the embedding representations of each word included in the corresponding word of each mention node; the initial features of the corresponding discourse node are obtained based on the embedding representations of each word included in the discourse sentence; based on the final output of the pre-trained language model, the hidden state corresponding to the start token position of the dialogue word sequence is used as the initial feature of the dialogue node.

[0037] Furthermore, in step A42, according to the following formula, the initial feature h (0) of each mention node is obtained based on the embedding representations of each word included in the corresponding word of each mention node:

[0038]

[0039] where w n is the embedding representation of the n-th word included in the corresponding word of this mention node, and k is the number of words included in the corresponding word of this mention node.

[0040] Furthermore, in step A42, according to the following steps, the initial feature z i of the corresponding discourse node is obtained based on the embedding representations of each word included in the discourse sentence:

[0041] For the i-th discourse sentence u i , first, based on the BiLSTM network, each word in the dialogue sentence u i is encoded to obtain the feature vectors i of each word included in the discourse sentence u

[0042] Then, the feature vectors i of each word included in the discourse sentence u are summed and averaged according to the number of words in the discourse sentence u i to obtain its mean feature The calculation formula is as follows:

[0043]

[0044] where t i represents the number of words in the i-th discourse sentence;

[0045] Finally, the sequential sequence composed of the mean features of each utterance sentence included in the dialogue is used as the input, and another parameter-independent BiLSTM network is used to encode it to obtain the initial features z of the corresponding discourse nodes of each utterance sentence included in the dialogue i ∈ {z1, z2,..., z m}, where m is the number of utterance sentences included in the dialogue

[0046] Furthermore, in step A5, based on the position information of the corresponding word of each mention node in the corresponding utterance sentence and the feature representation of the neighbor nodes of each mention node except the dialogue node, the feature representation of each mention node in the heterogeneous mention dialogue graph is updated based on the graph attention network; based on the position information of the corresponding utterance sentence of each discourse node in the dialogue and the feature representation of the neighbor nodes of each discourse node except the dialogue node, the feature representation of each discourse node in the heterogeneous mention dialogue graph is updated, including

[0047] A51. For the node n to be updated i , calculate its relative position information with its neighbor node n j according to the following formula

[0048] d u (n i , n j ) = uid(n i ) - uid(n j )

[0049] d p (n i , n j ) = p i - p j

[0050] where n j represents the neighbor node of node n i , If the node is the discourse node n u , then, uid(n u ) represents the position information of this utterance sentence in the dialogue, that is, the ID of this utterance sentence, and p u represents the sequential position of the first word of this utterance sentence in the dialogue. For the dialogue node n0 of the first utterance sentence of the dialogue, its uid(n0) = p0 = 0; if the node is the mention node n m , then, uid(n m ) represents the ID of the utterance sentence where the corresponding word of this mention node is located, and p m represents the position of the first word of the corresponding word of this mention node in the utterance sentence

[0051] A52. Encode through a non - linear function based on sine - position embedding, and calculate to obtain the node n to be updated i and its neighbor node n j 's relative - position encoding:

[0052]

[0053] where, I i,j represents the relative - position encoding between node n i and its neighbor node n j , W N and b N are both training parameters, PE(·) represents sine - position embedding, represents concatenation, and ReLU is an activation function;

[0054] A53. Based on the relative - position encoding I i,j , adopt an attention mechanism to update the feature representation of node n i :

[0055] First, calculate the attention - weight coefficient α i of node n j to its neighbor node n i,j according to the following formula:

[0056]

[0057]

[0058] where, represents the feature representation of node n i at the (l - 1) - th layer, represents the feature representation of neighbor node n j at the l - th layer; is a training parameter, and α i,j represents the attention - weight coefficient of node n i to its neighbor node n j ;

[0059] Then, based on the attention - weight coefficient α i of node n j to its neighbor node n i,j and the feature representation of its neighbor node n j , update its feature representation:

[0060]

[0061] where, and are both training parameters, ReLU is an activation function, respectively represent node ni Feature representations of the (l + 1)-th layer and the l-th layer Denote the neighbor node n j Feature representation at the l-th layer

[0062] Furthermore, in step A5, it further includes:

[0063] A54. Based on the fusion gate mechanism, a trade-off is made between the initial feature representation and the final feature representation output by the attention mechanism to obtain the final feature representation of the node n to be updated i where β

[0064]

[0065] is a gating coefficient vector and is calculated according to the following formula: i where ⊙ represents element-wise multiplication, that is, multiplying the elements at the corresponding positions, σ represents the activation function, W

[0066]

[0067]

[0068] and b b and b b are both training parameters

[0069] Furthermore, in step A6, based on the feature representations of the mentioned nodes and the utterance nodes obtained by updating, according to the correspondence between the mentioned nodes and the entities, the entity nodes and their feature representations are merged, and according to the correspondence between the utterance nodes and the speakers, the speaker nodes and their feature representations are merged; retain the edges in the heterogeneous mention dialogue graph except for the edges eliminated due to the merged nodes, and construct the feature representations of the edges based on the feature representations of the nodes at both ends of the edges, including:

[0070] The node merging calculation is performed according to the following formula:

[0071]

[0072] where e k represents the feature representation of the merged node, k is the node number, represents the number of nodes to be merged corresponding to the node e k and represents the feature representation of the t-th node among the nodes to be merged corresponding to the node e k ;

[0073] The feature representation of the edge is calculated according to the following formula:

[0074] eij = σ(W[e i ; e j +b)

[0075] where W and b are training parameters; σ is an activation function, e i , e j represent the feature representations of the i-th and j-th nodes of the entity dialogue graph respectively, and e ij represents the feature representation of the entity dialogue graph node e i and the edge between node e j .

[0076] Furthermore, in step A7, the path information between entity pairs in the entity dialogue graph is fused and reasoning is performed to obtain the relationship between entity pairs, including:

[0077] A71. Based on the attention mechanism, fuse each path between the corresponding entity nodes of the entity pair (e a , e b ):

[0078] First, calculate the attention weight coefficient α a , e b between each path between the corresponding entity nodes of the entity pair (e t ) according to the following formula:

[0079]

[0080]

[0081] where e a , e b represent the feature representations of the head entity e a and the tail entity e b of the entity pair respectively, is the feature representation of the t-th path from the head entity e a to the tail entity e b , σ represents the activation function, W l is a training parameter; α t is the attention weight, represents the s t -th power of the natural constant e, and k represents the number of paths between the entity pair (e a , e b );

[0082] Then, based on the attention weight coefficient α t and the feature representation of the path perform fusion according to the following formula:

[0083]

[0084] Among them, p a,b is the fused path feature representation;

[0085] A72. Based on the fused path feature representation p a,b , the feature representation u of the dialogue node dia and the feature representations of the head entity and the tail entity are concatenated to obtain the final entity pair feature representation:

[0086] H a,b = [e a ; e b ; |e a - e b |; e a ⊙ e b ; u dia ; p a,b

[0087] Among them, e a , e b respectively represent the feature representations of the head entity e a and the tail entity e b of the entity pair; u dia represents the feature representation of the dialogue node in the heterogeneous mention graph;

[0088] A73. Through a feed-forward neural network, for each predefined relationship r in the entity relation extraction task, a binary sigmoid classifier is used to determine the probability distribution P of each relationship r in the given entity pair (e a , e b ):

[0089] P(r|e a , e b ) = sigmoid(W1σ(W2H a,b + b1)+ b2)

[0090] Among them, W1, W2, b1, b2 are training parameters, and H a,b is the final entity pair feature representation, and σ is the activation function.

[0091] Furthermore, in step A71, only the one-hop paths between the entity nodes corresponding to the entity pair (e a , e b ) are fused. The one-hop path is also the path from the head entity e a to the tail entity e b that only passes through the intermediate node e o as the connection point, and the feature representation of its path can be expressed as:

[0092]

[0093] ​Among them, e ao represents the feature representation of the edge from entity e a to node e o .

[0094] Furthermore, in step A, an auxiliary task is further included. The auxiliary task includes a trigger word prediction task and a speaker prediction task; and according to the heterogeneous mention dialogue graph, the trigger word corresponding to the relationship is predicted through the trigger word prediction task, and the speaker corresponding to the relationship is predicted through the speaker prediction task;

[0095] In step A8, the dialogue relationship extraction model is jointly trained based on the cross-entropy loss of the trigger word prediction of the auxiliary task, the cross-entropy loss of the speaker prediction, and the binary cross-entropy loss of the relationship extraction.

[0096] Furthermore, the trigger word prediction task includes:

[0097] First, each predefined relationship r in the entity relationship extraction task is respectively mapped to a distributed embedding where e r is a One-Hot vector, which is 1 only at the position corresponding to the relationship and 0 at other positions;

[0098] Then, it is concatenated with each hidden feature output of the last layer of the pre-trained language model in step A41 to obtain

[0099]

[0100] Next, the probability distribution of the predicted label of the token word x i is calculated according to the following formula:

[0101] P(label = l|x i ) = softmax(W l z i + b l )

[0102] where and are training parameters, d r is the length of e r , that is, the number of relationships, d x is the length of the output feature vector of the pre-trained model, is the label set, representing the start word of the trigger word, the middle or end word of the trigger word, and the non-trigger word respectively.

[0103] Furthermore, the speaker prediction task includes:

[0104] In the heterogeneous mention dialogue graph, randomly select the speaker tags in the dialogue with a probability of 10%, and replace them with the special tag [MASK]. Then, input the modified sequence into the pre-trained language model to obtain the hidden states of each speaker's special tag, denoted as Then, calculate the probability distribution of the speaker through the following formula:

[0105]

[0106] where and are training parameters, d x is the length of the output feature vector of the pre-trained model, and Sp is the number of speakers.

[0107] Furthermore, in step A8, based on the cross-entropy loss of trigger word prediction, the cross-entropy loss of speaker prediction, and the binary cross-entropy loss of relation extraction in the auxiliary task, jointly train the dialogue relation extraction model. The loss function for joint training is:

[0108]

[0109] where is the binary cross-entropy loss of relation extraction, is the cross-entropy loss of trigger word prediction, is the cross-entropy loss of speaker prediction, and α is the weight of the auxiliary task;

[0110] where is calculated according to the following formula:

[0111]

[0112] where N1 is the number of entity pairs to be predicted, p n is the vector composed of the predicted probability distributions of each relation corresponding to the entity pair; y n ∈(0, 1) is the vector composed of the true labels of each relation corresponding to the entity pair;

[0113] is calculated according to the following formula:

[0114]

[0115] where length is the length of the dialogue, i.e., the number of words, is the vector composed of the true trigger word labels, is the vector composed of the predicted trigger word label probability distributions;

[0116]

[0117] Wherein, N2 is the number of discourse sentences, and M is the number of speakers in the conversation. is a vector composed of real speaker labels. is a vector composed of predicted speaker probability distributions.

[0118] The beneficial effects of the present invention are as follows:

[0119] The present invention constructs a heterogeneous mention dialogue graph with a smaller granularity, and introduces relative position awareness information for node features with similar structures in the heterogeneous mention dialogue graph, helping the model extract more discriminative mention node features, thereby improving the accuracy of relation extraction; during the training process, two auxiliary tasks, namely speaker prediction and trigger word prediction tasks, are further introduced, and by capturing the unique features of the speaker and discovering highly supportive information for the correct relationship, the performance of dialogue-based relation extraction is enhanced. Description of the Drawings

[0120] Figure 1 is a training process diagram of the dialogue relation extraction model in the embodiment of the present invention;

[0121] Figure 2 is a reference instance diagram of the dataset used in the model training in the embodiment of the present invention. Detailed Embodiments

[0122] The present invention aims to propose a dialogue relation extraction method that integrates position awareness refinement, solves the interference caused by similar structural information in existing dialogue texts, cannot classify nodes, and has the problem of low relation extraction accuracy due to the failure to consider the uniqueness of dialogue data. During the model training process of the present invention, first, the dialogue for training and the entity pairs for which relations need to be extracted are input, and the dependency syntactic analysis is performed on the dialogue text; then, a heterogeneous mention dialogue graph is constructed according to the word segmentation and syntactic analysis results; next, the dialogue relation extraction task is carried out: by using the position awareness refinement graph attention network on the heterogeneous mention dialogue graph, the updated node features are obtained; then, an entity dialogue graph is constructed based on the updated node features, and then, the path information between the entity pairs in the entity dialogue graph is fused, and the relation between the entity pairs is inferred through reasoning; in addition, to improve the relation extraction performance, the present invention also introduces auxiliary tasks during the training process, and captures the unique features of the speaker and discovers highly supportive information for the correct relationship through the auxiliary tasks; finally, through the joint training of the relation extraction task, the auxiliary tasks and based on the multi-task learning framework, a dialogue relation extraction model is obtained.

[0123] Embodiment:

[0124] The method for refining dialogue relationship extraction with fused location awareness in this embodiment includes two major parts: training a dialogue relationship extraction model and performing a dialogue relationship extraction task using the trained model. Among them, the training process of the dialogue relationship extraction model is as Figure 1 shown and includes the following steps:

[0125] S1. Data preprocessing:

[0126] In this step, the input data is preprocessed. The input data includes the dialogue for training and the entity pairs for which the relationship needs to be extracted. Among them, the dialogue for training is a dialogue dataset in text form. We use natural language analysis tools to tokenize the original text in the dialogue dataset and perform dependency syntactic analysis to obtain the tokenization result.

[0127] Specifically, the dialogue dataset adopts the publicly available dialogue relationship extraction dataset DialogRE based on manual annotation, which contains 1,788 dialogues U. For each dialogue U, it contains several utterance sentences T, and an utterance sentence is defined when the speaker changes. Each utterance sentence is further divided into several clauses and several words. Specific dataset examples are as Figure 2 shown.

[0128] The specific preprocessing process is as follows: First, encode the dialogue content source file into UTF-8 format and remove non-text parts such as XML tags inside. Then use the StanfordCoreNLP toolkit provided by Stanford to perform syntactic component analysis on the text, thereby obtaining a dependency syntactic analysis tree. The dependency syntactic analysis tree is the structured knowledge that can describe the dependency relationship between words in one or more sentences.

[0129] According to the obtained dependency syntactic analysis result, the words with the same text as the entities in the entity pair and the words with a referential relationship are used as mention words. For example, for the entity "University of Electronic Science and Technology", mentions such as "UESTC", "Chengdu University of Electronic Science and Technology", and "UESTC" in the text all point to this entity. Therefore, "UESTC", "Chengdu University of Electronic Science and Technology", and "UESTC" can all be used as mention words; and based on the dialogue, according to the change of the speaker or the pause in the dialogue, the utterance sentences are segmented and obtained.

[0130] S2. Construct a heterogeneous mention dialogue graph and initialize node features:

[0131] In this step, the dependency syntactic analysis tree is input into the BERT model to output the feature representation containing different granularity semantic information in the dialogue, thereby constructing a heterogeneous mention dialogue graph; and the node features in the heterogeneous mention dialogue graph are initialized.

[0132] In this embodiment, three types of nodes and four types of edges are set in the heterogeneous mention dialogue graph, including:

[0133] Mention Nodes: Each mention word is set as a mention node, and mention words with the same text but different positions in the conversation are set as different mention nodes; different positions in the conversation include that mention words with the same text appear in different utterance sentences and that mention words with the same text appear in different positions within the same utterance sentence;

[0134] Utterance Nodes: Each utterance sentence is set as an utterance node;

[0135] Conversation Node: The conversation is set as a conversation node;

[0136] The edges formed by pairwise connections between each pair of mention nodes corresponding to the same entity constitute mention dependency edges;

[0137] The edges formed by pairwise connections between each pair of utterance nodes corresponding to the same speaker, and the edges formed by pairwise connections between each pair of mention nodes corresponding to the same speaker and different entities, constitute speaker dependency edges;

[0138] The edges formed by the edges between a mention node and the utterance node corresponding to the utterance sentence to which the word corresponding to the mention node belongs constitute utterance dependency edges;

[0139] The edges formed by the edges between a mention node and the conversation node constitute conversation dependency edges;

[0140] After setting the nodes and edges of the heterogeneous mention-conversation graph, obtain the initial vector representation of the nodes. The specific node initialization includes the following steps:

[0141] (1) For each utterance sentence in the conversation, first, concatenate the speaker label and the sequence formed by the words of the utterance sentence in sequence to form an utterance sentence word sequence, and then, concatenate the start label of the conversation, the sequence formed by each utterance sentence word sequence in sequence, and the end label of the conversation to form a conversation word sequence; input the obtained conversation word sequence into a pre-trained language model to obtain an initial word embedding representation;

[0142] Then, according to the obtained initial word embedding representation, for each utterance sentence, respectively concatenate the embedding representation i of its speaker s with the initial embedding representation ij of each of its words x to obtain the embedding representation w ij of each of its words:

[0143]

[0144] where represents the concatenation operation; w ij represents the embedding representation of the j-th word in the utterance sentence u i ;

[0145] (2) Based on the embedding representations of each word included in the corresponding word of each mention node, the initial feature h of each mention node (0) is calculated as follows:

[0146]

[0147] where w n is the embedding representation of the nth word included in the corresponding word of this mention node, and k is the number of words included in the corresponding word of this mention node.

[0148] (3) Based on the embedding representations of each word included in the discourse sentence, the initial feature z of its corresponding discourse node is obtained i , and the process is as follows:

[0149] For the ith discourse sentence u i , first, based on the BiLSTM network, each word in the dialogue sentence u i is encoded to obtain the feature vectors of each word included in the discourse sentence u i

[0150]

[0151]

[0152]

[0153] where respectively represent the jth and j+1th hidden feature representations obtained by forward encoding of BiLSTM in the discourse sentence u i ; and respectively represent the jth and j-1th hidden feature representations obtained by backward encoding of BiLSTM in the discourse sentence u i ; LSTM b , LSTM f represent the feature extraction in the forward and backward directions of the BiLSTM encoder;

[0154] Then, the feature vectors i of each word included in the discourse sentence u are summed up and averaged according to the number of words in the discourse sentence u i to obtain its mean feature The calculation formula is as follows:

[0155]

[0156] where t i represents the number of words in the ith discourse sentence;

[0157] Finally, the sequential sequence formed by the mean features of each utterance sentence included in the dialogue is used as the input, and another parameter-independent BiLSTM network is used to encode it to obtain the initial features z of the corresponding discourse nodes of each utterance sentence included in the dialogue i ∈ {z1, z2,..., z m}, where m is the number of utterance sentences included in the dialogue

[0158] (4) Use BERT to encode the entire dialogue, and use the hidden state representation at the [CLS] position corresponding to the last layer of the BERT vocabulary as the initial dialogue node representation u dia .

[0159] The core of the present invention lies in obtaining features at a finer granularity. Therefore, in addition to the above methods, other methods in the prior art can also be used for the initialization of the node features of the present invention

[0160] S3. Parallelly execute the relationship prediction task and the auxiliary task

[0161] Generally, a relationship triple is composed of (head entity, relationship, tail entity). The relationship prediction task is to given the head entity and the tail entity included in the dialogue, and extract the relationship between the head entity and the tail entity according to the dialogue content. The auxiliary task is to construct a speaker prediction and a trigger word prediction task to capture the unique features of the speaker and discover highly supportive information for the correct relationship, thereby improving the accuracy of relationship prediction

[0162] Specifically, for the relationship prediction task, in this embodiment, first, based on the graph attention network, update the node features in the heterogeneous mention dialogue graph, and then, based on the updated node features, construct an entity dialogue graph, fuse the path information between entity pairs in the entity dialogue graph, and perform reasoning to obtain the relationship between entity pairs, including two sub-processes: node feature update and relationship prediction

[0163] I. Node feature update

[0164] According to the position information of the corresponding word of each mention node in the corresponding utterance sentence, and the feature representations of the neighbor nodes of each mention node other than the dialogue node, based on the graph attention network, update the feature representations of each mention node in the heterogeneous mention dialogue graph; according to the position information of the corresponding utterance sentence of each discourse node in the dialogue, and the feature representations of the neighbor nodes of each discourse node other than the dialogue node, based on the graph attention network, update the feature representations of each discourse node in the heterogeneous mention dialogue graph; the detailed implementation process is introduced as follows

[0165] (1) For the node n to be updated i , calculate its neighbor node n according to the following formulaj Relative position information:

[0166] d u (n i ,n j ) = uid(n i ) - uid(n j )

[0167] d p (n i ,n j ) = p i - p j

[0168] Where n j represents the neighbor node of node n i . If the node is the utterance node n , then, uid(n u ) represents the position information of this utterance sentence in the dialogue, that is, the ID of this utterance sentence, and p u represents the sequential position of the first word of this utterance sentence in the dialogue. For the dialogue node n0 of the first utterance sentence of the dialogue, its uid(n0) = p0 = 0; if the node is the mention node n u , then, uid(n m ) represents the ID of the utterance sentence where the word corresponding to this mention node is located, and p m represents the position of the first word of the word corresponding to this mention node in the utterance sentence; m

[0169] (2) Encode through a non - linear function based on sine - position embedding, and calculate the relative - position encoding between the node n i to be updated and its neighbor node n j :

[0170]

[0171] Where I i,j represents the relative - position encoding between node n i and its neighbor node n j , W N and b N are both training parameters, PE(·) represents sine - position embedding, represents concatenation, and ReLU is the activation function;

[0172] (3) Based on the relative - position encoding I i,j , adopt the attention mechanism to update the feature representation of node n i :

[0173] First, calculate the node n iThe attention weight coefficient α of its neighbor node n j is as follows: i,j :

[0174]

[0175]

[0176] wherein,

[0177] represents the feature representation of node n i at the (l - 1)-th layer, represents the feature representation of neighbor node n j at the l-th layer; is a training parameter, and α i,j represents the attention weight coefficient of node n i to its neighbor node n j ;

[0178] Then, based on the attention weight coefficient α of node n i to its neighbor node n j and the feature representation of its neighbor node n i,j and the feature representation of its neighbor node n j is updated as follows:

[0179]

[0180] wherein, and are both training parameters, ReLU is an activation function, respectively represent the feature representations of node n i at the (l + 1)-th layer and the l-th layer, represents the feature representation of neighbor node n j at the l-th layer.

[0181] (4) Based on the fusion gate mechanism, a trade-off is made between the initial feature representation and the final feature representation output by the attention mechanism to obtain the final feature representation i of the node n to be updated

[0182]

[0183] wherein, β i is a gating coefficient vector and is calculated according to the following formula:

[0184]

[0185]

[0186] Among them, ⊙ represents element-wise multiplication, that is, multiplying elements at corresponding positions, σ represents the activation function, W b and b b are both training parameters.

[0187] II. Relationship Prediction:

[0188] Based on the feature representations of the mention nodes and utterance nodes obtained by updating, entity nodes and their feature representations are obtained by merging according to the corresponding relationships between the mention nodes and entities, and speaker nodes and their feature representations are obtained by merging according to the corresponding relationships between the utterance nodes and speakers; in the heterogeneous mention-dialogue graph, edges other than those eliminated due to node merging are retained, and the feature representations of the edges are constructed based on the feature representations of the nodes at both ends of the edges, thereby constructing an entity-dialogue graph. Then, the path information between entity pairs in the entity-dialogue graph is fused, and reasoning is performed to obtain the relationship between entity pairs. The detailed implementation process is introduced as follows:

[0189] (1) The merging calculation of nodes is performed according to the following formula:

[0190]

[0191] Among them, e k represents the feature representation of the merged node, k is the node serial number, represents the number of nodes to be merged corresponding to node e k , represents the feature representation of the t-th node among the nodes to be merged corresponding to node e k .

[0192] (2) The feature representation of the edge is calculated according to the following formula:

[0193] e ij =σ(W[e i ; e j )+b

[0194] Among them, W and b are training parameters; σ is the activation function, e i , e j respectively represent the feature representations of the i-th and j-th nodes of the entity-dialogue graph, and e ij represents the feature representation of the edge between node e i and node e j in the entity-dialogue graph.

[0195] (3) Based on the attention mechanism, the paths between the entity nodes corresponding to the entity pair (e a , e b ) are fused:

[0196] For the entity pair (e a , eb ) There exists a certain path in the entity dialogue graph. Taking the intermediate node e in the path o as the connection point, the head entity e a to the tail entity e b of the t-th path can be expressed as:

[0197]

[0198] where, e ao represents the edge feature from entity e a to node e o ;

[0199] Calculate the attention weight coefficient α a between each path of the corresponding entity nodes of the entity pair (e b , e t ):

[0200]

[0201]

[0202] where, e a , e b respectively represent the feature representations of the head entity e a and the tail entity e b of the entity pair, is the feature representation of the t-th path from the head entity e a to the tail entity e b , σ represents the activation function, W l is the training parameter; α t is the attention weight, represents the s t power of the natural constant e, k represents the number of paths between the entity pair (e a , e b );

[0203] Then, based on the attention weight coefficient α t and the feature representation of the path, perform fusion according to the following formula:

[0204]

[0205] where, p a,b is the fused path feature representation.

[0206] (4) Based on the fused path feature representation p a,b , the feature representation u diaAnd concatenate the feature representations of the head entity and the tail entity to obtain the final entity pair feature representation:

[0207] H a,b = [e a ; e b ; |e a -e b |; e a ⊙e b ; u dia ; p a,b

[0208] Among them, e a , e b respectively represent the feature representations of the head entity e a and the tail entity e b of the entity pair; u dia represents the feature representation of the dialogue node in the heterogeneous mention graph;

[0209] (5) Through a feed-forward neural network, for each predefined relationship r in the entity relation extraction task, a binary sigmoid classifier is adopted to determine the probability distribution P of each relationship r in the given entity pair (e a , e b ):

[0210] P(r|e a , e b ) = sigmoid(W1σ(W2H a,b +b1)+b2)

[0211] Among them, W1, W2, b1, and b2 are training parameters, and H a,b is the final entity pair feature representation, and σ is the activation function.

[0212] For the auxiliary tasks, in this embodiment, by constructing a trigger word prediction task and a speaker prediction task, according to the heterogeneous mention dialogue graph, the trigger word corresponding to the relationship is predicted through the trigger word prediction task, and the speaker corresponding to the relationship is predicted through the speaker prediction task. Specifically:

[0213] (1) Trigger word prediction task:

[0214] First, map each predefined relationship r in the entity relation extraction task to a distributed embedding respectively. e r is a One-Hot vector, which is 1 only at the position corresponding to the relationship and 0 at other positions;

[0215] Then, connect it with each hidden feature output of the last layer of the pre-trained language model to be

[0216] ​

[0217] Then calculate the probability distribution of the predicted label of the marked word x according to the following formula: i :

[0218] P(label=l|x i )=softmax(W l z i +b l )

[0219] Wherein, and are training parameters, d r is e r 's length, i.e., the number of relationships, d x is the length of the output feature vector of the pre-trained model, is the label set, representing the trigger word start word, the trigger word middle or end word, and the non-trigger word respectively.

[0220] (2) Speaker prediction task:

[0221] In the heterogeneous mention dialogue graph, randomly select the speaker label in the dialogue with a probability of 10%, and replace it with the special token [MASK]. Then, input the modified sequence into the pre-trained language model to obtain the hidden state of each speaker special token, denoted as Then, calculate the probability distribution of the speaker according to the following formula:

[0222]

[0223] Wherein, and are training parameters, d x is the length of the output feature vector of the pre-trained model, and Sp is the number of speakers.

[0224] S4. Multi-task joint learning:

[0225] In this step, the relationship prediction task, the trigger word prediction task, and the speaker prediction task share the BERT encoder and are jointly trained based on the multi-task learning framework. The loss function for joint training is:

[0226]

[0227] Wherein, is the binary cross-entropy loss for relationship extraction, is the cross-entropy loss for trigger word prediction, is the cross-entropy loss for speaker prediction, and α is the weight of the auxiliary task;

[0228] Wherein, is calculated according to the following formula:

[0229]

[0230] Among them, N1 is the number of entity pairs to be predicted, and p n is a vector composed of the predicted probability distributions of each relationship corresponding to the entity pair; y n ∈(0, 1) is a vector composed of the true labels of each relationship corresponding to the entity pair;

[0231] It is calculated according to the following formula:

[0232]

[0233] Among them, length is the dialogue length, that is, the number of words, is a vector composed of the true trigger word labels, is a vector composed of the predicted trigger word label probability distributions;

[0234]

[0235] Among them, N2 is the number of utterance sentences, M is the number of speakers in the dialogue, is a vector composed of the true speaker labels, is a vector composed of the predicted speaker probability distributions.

[0236] Based on the above process, a dialogue relationship extraction model can be established. In practical applications, taking the dialogue text and the entity pair containing the head and tail entities as input, and using the trained dialogue relationship, the extraction model can obtain the relationship between the head and tail entities.

[0237] Although the present invention has been described herein with reference to the embodiments of the present invention, the above embodiments are only preferred embodiments of the present invention, and the embodiments of the present invention are not limited by the above embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, and these modifications and embodiments will fall within the scope of the principles and spirit disclosed in this application.

Claims

1. A method for extracting dialogue relationships that integrates position-aware refinement, characterized in that, It includes the following steps: A. Training a dialogue relation extraction model: A1. Inputting the dialogue for training and the entity pairs for which relations need to be extracted; A2. Tokenizing the original text of the dialogue and performing dependency syntactic analysis. Then, according to the obtained dependency syntactic analysis results, the words with the same text as the entities in the entity pair and the words with a referential relationship are used as mention words; Based on the dialogue, segmenting to obtain utterance sentences according to the change of the speaker or the pauses in the dialogue; A3. Constructing a heterogeneous mention dialogue graph: The nodes of the heterogeneous mention dialogue graph include mention nodes, utterance nodes, and dialogue nodes, and the edges include mention dependency edges, speaker dependency edges, utterance dependency edges, and dialogue dependency edges, and are constructed according to the following definitions: Mention nodes: Each mention word is set as a mention node, and mention words with the same text but different positions in the dialogue are set as different mention nodes. Different positions in the dialogue include that mention words with the same text appear in different utterance sentences and that mention words with the same text appear in different positions in the same utterance sentence; Utterance nodes: Each utterance sentence is respectively set as an utterance node; Dialogue nodes: The dialogue is set as a dialogue node; The edges formed by connecting pairwise the mention nodes corresponding to the same entity constitute mention dependency edges; The edges formed by connecting pairwise the utterance nodes corresponding to the same speaker, and the edges formed by connecting pairwise the mention nodes corresponding to the same speaker and different entities constitute speaker dependency edges; The edges between the mention node and the utterance node corresponding to the utterance sentence to which the corresponding word of the mention node belongs constitute utterance dependency edges; The edges between the mention node and the dialogue node constitute dialogue dependency edges; A4. Initializing the node features in the heterogeneous mention dialogue graph; A5. Based on the position information of the corresponding words of each mention node in the corresponding utterance sentence, and the feature representations of the neighbor nodes of each mention node except the dialogue node, updating the feature representations of each mention node in the heterogeneous mention dialogue graph based on the graph attention network; based on the position information of the utterance sentences corresponding to each utterance node in the dialogue, and the feature representations of the neighbor nodes of each utterance node except the dialogue node, updating the feature representations of each utterance node in the heterogeneous mention dialogue graph based on the graph attention network; A6. Constructing an entity dialogue graph: Based on the updated feature representations of the mention nodes and utterance nodes, merging to obtain entity nodes and their feature representations according to the corresponding relationship between the mention nodes and the entities, and merging to obtain speaker nodes and their feature representations according to the corresponding relationship between the utterance nodes and the speakers; Retaining the edges in the heterogeneous mention dialogue graph except for the edges eliminated due to node merging, and constructing the feature representations of the edges based on the feature representations of the nodes at both ends of the edges; A7. Fusing the path information between entity pairs in the entity dialogue graph and performing reasoning to obtain the relationship between entity pairs; A8. Repeatedly executing steps A1 - A7, training the dialogue relation extraction model based on the binary cross - entropy loss of relation extraction until the model converges or reaches the set number of training rounds; B. Performing a dialogue relation extraction task: Taking the dialogue text and the entity pair of the relationship to be extracted as input, and using the trained dialogue relationship extraction model to obtain the relationship between the entities of the entity pair.

2. A dialogue relationship extraction method integrating position perception refinement as described in claim 1, characterized in that In step A2, the original text of the dialogue is tokenized and dependency syntactic analysis is performed, specifically including: First, remove the non-text part of the original text and perform format encoding; then, use natural language processing tools to tokenize and perform dependency syntactic analysis on the text to obtain a dependency syntactic analysis tree.

3. A dialogue relationship extraction method integrating position perception refinement as described in claim 1, characterized in that In step A4, the node features in the heterogeneous mention dialogue graph are initialized, including: A41. For each utterance sentence of the dialogue, first, concatenate the speaker tag and the sequence composed of the words of the utterance sentence in sequence to form an utterance sentence word sequence, and then, concatenate the start tag of the dialogue, the sequence composed of each utterance sentence word sequence in sequence, and the end tag of the dialogue to form a dialogue word sequence; input the obtained dialogue word sequence into a pre-trained language model to obtain an initial word embedding representation. Then, according to the obtained initial word embedding representations, for each utterance sentence, the embedding representation of its speaker s i is concatenated with the initial embedding representation of each of its words x ij to obtain the embedding representation w of each of its words : ij ​ Among them, represents a splicing operation; w ij represents the embedding representation of the j-th word in the utterance sentence u i ; A42. The embedding representation w of each word obtained according to step A41 ij , based on the embedding representations of the words included in the corresponding words of each mention node, obtain the initial features of each mention node; based on the embedding representations of the words included in the discourse sentence, obtain the initial features of its corresponding discourse node; based on the final output of the pre-trained language model, use the hidden state corresponding to the start token position of the dialogue word sequence as the initial feature of the dialogue node.

4. A dialogue relationship extraction method integrating position perception refinement as described in claim 3, characterized in that In step A42, according to the following formula, based on the embedding representations of each word included in the corresponding word of each mentioned node, the initial feature h of each mentioned node is obtained (0) : where w n is the embedding representation of the n-th word contained in the corresponding word of the mentioned node, and k is the number of words contained in the corresponding word of the mentioned node.

5. A dialogue relationship extraction method integrating position perception refinement as described in claim 3, characterized in that In step A42, based on the embedding representations of the words included in the utterance sentence, the initial feature z of its corresponding utterance node is obtained according to the following steps i : For the i-th utterance u i , first, based on the BiLSTM network, each word in the dialogue utterance u i is encoded to obtain the feature vectors of the words included in the utterance u i ​ Then, for the utterance sentence u i the feature vectors of each word it contains are summed up and averaged according to the number of words in the utterance sentence u i to obtain its mean feature The calculation formula is as follows: where t i represents the number of words in the i-th utterance; Finally, use the sequential sequence composed of the mean features of each utterance sentence included in the dialogue as the input, and use another parameter-independent BiLSTM network to encode it to obtain the initial features z of the dialogue nodes corresponding to each utterance sentence included in the dialogue i ∈{z1, z2, …, z m}, where m is the number of utterance sentences included in the dialogue.

6. A method for extracting dialogue relationships with refined fusion of location awareness as described in any one of claims 1, 2, or 3, characterized in that In step A5, according to the position information of the corresponding words of each mention node in the corresponding utterance sentence, and the feature representations of the neighbor nodes of each mention node except the dialogue node, based on the graph attention network, update the feature representations of each mention node in the heterogeneous mention dialogue graph; according to the position information of the corresponding utterance sentence of each dialogue node in the dialogue, and the feature representations of the neighbor nodes of each dialogue node except the dialogue node, based on the graph attention network, update the feature representations of each dialogue node in the heterogeneous mention dialogue graph, including: A51. For the node n to be updated i , calculate its relative position information with its neighbor node n j according to the following formula: d u (n i ,n j ) = uid(n i ) - uid(n j ) d p (n i ,n j )=p i -p j where n j represents the neighbor nodes of node n i . If the node is the utterance node n u , then, uid(n u ) represents the position information of the utterance sentence in the dialogue, that is, the ID of the utterance sentence, p u represents the sequential position of the first word of the utterance sentence in the dialogue. For the dialogue node n0 of the first utterance sentence of the dialogue, its uid(n0) = p0 = 0; if the node is the mention node n m , then, uid(n m ) represents the ID of the utterance sentence corresponding to the word of the mention node, p m represents the position of the first word of the word corresponding to the mention node in the utterance sentence; A52. Calculate the relative position encoding of the node n to be updated and its neighbor node n by encoding through a non - linear function based on sine position embedding: i and its neighbor node n j : Among them, I i,j represents the relative position encoding of node n i with its neighbor node n j W N and b N are both training parameters, PE(·) represents the sine position embedding, denotes concatenation, and ReLU is the activation function; A53. Relative Position Encoding I i,j , using the attention mechanism, update the feature representation of node n i : First, calculate the node n according to the following formula i for its neighbor node n j the attention weight coefficient α i,j : Among them, represents the feature representation of node n i at the (l - 1)-th layer, represents the feature representation of neighbor node n j at the l-th layer; is a training parameter, α i,j represents the attention weight coefficient of node n i for its neighbor node n j ; Then, based on node n i the attention weight coefficient α j of its neighbor node n i,j and the feature representation of its neighbor node n j are used to update its feature representation: Among them, and are both training parameters, and ReLU is an activation function. respectively represent the feature representations of node n i at the (l + 1)-th layer and the l-th layer. represents the feature representation of neighbor node n j at the l-th layer.

7. A dialogue relationship extraction method integrating position perception refinement as described in claim 6, characterized in that In step A5, it further includes: A54. Based on the fusion gate mechanism, a trade-off is made between the initial feature representation and the final feature representation output by the attention mechanism to obtain the node n to be updated i The final feature representation where β i is the gating coefficient vector and is calculated according to the following formula: Among them, ⊙ represents element-wise multiplication, that is, multiplying elements at corresponding positions, σ represents the activation function, W b and b b are both training parameters.

8. A method for extracting dialogue relationships by fusing location perception refinement according to any one of claims 1, 2, or 3, characterized in that In step A6, based on the updated feature representations of the mention nodes and dialogue nodes, merge to obtain entity nodes and their feature representations according to the corresponding relationship between the mention nodes and the entities, and merge to obtain speaker nodes and their feature representations according to the corresponding relationship between the dialogue nodes and the speakers; Retain the edges in the heterogeneous mention dialogue graph except for the edges eliminated due to the merged nodes, and construct the feature representations of the edges based on the feature representations of the nodes at both ends of the edges, including: Perform the merging calculation of the nodes according to the following formula: Among them, e k represents the node feature representation obtained after merging, k is the node serial number, represents node e k the number of nodes to be merged corresponding to it, represents node e k the feature representation of the t-th node among the nodes to be merged corresponding to it; Calculate the feature representation of the edge according to the following formula: e ij = σ(W[e i ; e j + b) Among them, W and b are training parameters; σ is the activation function, e i , e j respectively represent the feature representations of the i-th and j-th nodes of the entity dialogue graph, e ij represents the entity dialogue graph node e i and the node e j the feature representation of the edge between them.

9. A method for extracting dialogue relationships with refined fusion of location awareness according to any one of claims 1, 2, or 3, characterized in that In step A7, fuse the path information between the entity pairs in the entity dialogue graph and perform reasoning to obtain the relationship between the entity pairs, including: A71. Based on the attention mechanism, fuse each path between the entity nodes corresponding to the entity pair (e a , e b ): First, calculate the attention weight coefficient α a between the entity nodes corresponding to the entity pair (e b ) for each path according to the following formula: t : Among them, e a and e b respectively represent the feature representations of the head entity e a and the tail entity e b . is the feature representation of the t-th path from the head entity e a to the tail entity e b . σ represents the activation function, and W l are the training parameters; α t is the attention weight, represents the s t -th power of the natural constant e, and k represents the number of paths between the entity pair (e a , e b ). Then, based on the attention weight coefficient α t and the feature representation of the path perform fusion according to the following formula: Among them, p a,b is the fused path feature representation; A72. Concatenate the path feature representation p after fusion a,b , the feature representation u of the dialogue node dia and the feature representations of the head entity and the tail entity to obtain the final entity pair feature representation: H a,b = [e a ; e b ; |e a -e b |; e a ⊙e b ; u dia ; p a,b ​ Among them, e a , e b respectively represent the feature representations of the head entity e a and the tail entity e b ; u dia represents the feature representation of the dialogue node in the heterogeneous mention graph; A73. Through a feedforward neural network, for each predefined relation r in the entity relation extraction task, a binary sigmoid classifier is used to determine the probability distribution P of each relation r in the given entity pair (e a , e b ): P(r|e a , e b ) = sigmoid(W1σ(W2H a,b + b1) + b2) Among them, W1, W2, b1, and b2 are training parameters, and H a,b is the final entity pair feature representation, and σ is the activation function.

10. A dialogue relationship extraction method integrating position perception refinement as described in claim 9, characterized in that In step A71, only the one-hop paths between the entity nodes corresponding to the entity pair (e a , e b ) are fused. The one-hop path is also the path from the head entity e a to the tail entity e b that only passes through the intermediate node e o as the connection point, and the feature representation of its path can be expressed as: Among them, e ao represents the feature representation of the edge from entity e a to node e o .

11. A dialogue relationship extraction method integrating position perception refinement as described in claim 3, characterized in that In step A, an auxiliary task is further included, and the auxiliary task includes a trigger word prediction task and a speaker prediction task; and according to the heterogeneous mention dialogue graph, the trigger word prediction corresponding to the relationship is performed through the trigger word prediction task, and the speaker prediction corresponding to the relationship is performed through the speaker prediction task; In step A8, based on the cross-entropy loss of the trigger word prediction of the auxiliary task, the cross-entropy loss of the speaker prediction, and the binary cross-entropy loss of the relationship extraction, the dialogue relationship extraction model is jointly trained.

12. A dialogue relationship extraction method incorporating position-aware refinement according to claim 11, wherein: The trigger word prediction task includes: First, each predefined relation r in the entity relation extraction task is mapped to a distributed embedding respectively in which, e r is a One-Hot vector, which is 1 only at the position corresponding to the relation and 0 at other positions; Then, it is connected to each hidden feature output of the last layer of the pre-trained language model in step A41 as Then calculate the probability distribution of the predicted label of the marked word x according to the following formula: i : P(label = l|x i ) = softmax(W l z i + b l ) Among them, and are training parameters, d r is the length of e r i.e., the number of relationships, d x is the length of the output feature vector of the pre-trained model, is a set of labels, representing the start word of the trigger word, the middle or end word of the trigger word, and the non-trigger word respectively.

13. A dialogue relationship extraction method incorporating position-aware refinement according to claim 9, wherein: The speaker prediction task includes: In the heterogeneous mention dialogue graph, randomly select the speaker token in the dialogue with a probability of 10%, and replace it with the special token [MASK]. Then, input the modified sequence into the pre-trained language model to obtain the hidden state of each speaker's special token, denoted as Then, calculate the probability distribution of the speaker through the following formula: Among them, and are training parameters, d x is the length of the output feature vector of the pre-trained model, and Sp is the number of speakers.

14. A method for extracting dialogue relationship with refined fusion of location awareness according to any one of claims 11, 12 or 13, characterized in that In step A8, based on the cross-entropy loss of the trigger word prediction of the auxiliary task, the cross-entropy loss of the speaker prediction, and the binary cross-entropy loss of the relationship extraction, the dialogue relationship extraction model is jointly trained, and the loss function of its joint training is: Among them, is the binary cross-entropy loss for relation extraction, is the cross-entropy loss for trigger word prediction, is the cross-entropy loss for speaker prediction, and α is the weight of the auxiliary task; Among them, is calculated according to the following formula: where N1 is the number of entity pairs to be predicted, and p n is a vector composed of the predicted probability distributions of each relationship corresponding to the entity pair; y n ∈(0, 1) is a vector composed of the true labels of each relationship corresponding to the entity pair; Calculated according to the following formula: where length is the dialogue length, i.e., the number of words, is a vector composed of real trigger word tags, is a vector composed of the predicted trigger word tag probability distribution; where N2 is the number of utterance sentences and M is the number of speakers in the conversation, is a vector composed of true speaker labels, is a vector composed of predicted speaker probability distributions.

Citation Information

Patent Citations

  • Medical consultation dialogue system and method applying heterogeneous graph neural network

    CN112271001A

  • Heterogeneous graph structure multi-dialogue sentiment analysis method based on topic semantic enhancement

    CN114911932A