Role and word representation learning method based on binary self-attention driving
By introducing dual self-attention-driven role and word representation learning methods in the information extraction technology, the problem of failure to fully explore the semantic relationship between roles and words in the prior art is solved, and more accurate role and word representation learning is achieved.
Patent Information
- Application Number
- CN202510562662.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The prior art fails to fully explore the relationship between the semantics of each role in the event type and the semantics of each word that acts as an argument in the chapter-level event argument extraction.
The role and word representation learning method based on dual self-attention drive is adopted. Through the feature fusion module, role-word self-attention module and word-role self-attention module, the interactive semantics between roles and words are deeply explored, and optimized through multi-layer perceptron and span predictors.
It realizes more accurately capturing role semantic information and word role semantic information in events, improving the accuracy and effectiveness of meta-role and word representations in the chapter.
Smart Images

Figure CN120067306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information extraction, and particularly to a method for learning the representation of roles and words driven by dual self-attention. Background Art
[0002] Event argument extraction at the discourse level is a key task in information extraction, which extracts event-related arguments from a discourse and accurately identifies their roles. The existing role-based span selection strategy mainly captures role representations and discourse representations through pre-trained language models and prompts, so as to construct a span selector to predict spans for each role. However, the existing methods do not fully explore the association between the semantics of each role in the event type and the semantics of each word serving as an argument in the discourse. Summary of the Invention
[0003] In view of the above situation, the main object of the present invention is to propose a method for learning the representation of roles and words driven by dual self-attention to solve the above technical problems.
[0004] The present invention proposes a method for learning the representation of roles and words driven by dual self-attention, and the method includes the following steps: Step 1: Construct a feature fusion module based on a linear transformation mechanism and a gated neural network, and construct a role-word self-attention module and a word-role self-attention module based on a bidirectional self-attention mechanism. The feature fusion module, the role-word self-attention module, the word-role self-attention module, a multi-layer perceptron, and a span predictor constitute a prediction model; Step 2: Input the text and give the discourse, perform syntactic and semantic analysis on the discourse, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings; Step 3: Use the feature fusion module to capture the span association between the pattern semantics of the role and the context semantics of the word for different semantic embeddings, and obtain the initial embedding representation of the role and the initial embedding representation of the word; Step 4: Based on the role-word self-attention module, use the bidirectional self-attention mechanism to calculate the attention weight of the role-word, and update the initial embedding representation of the role to obtain an updated role embedding representation; Based on the word-role self-attention module, use the bidirectional self-attention mechanism to calculate the attention weight of the word-role, and update the initial embedding representation of the word to obtain an updated word embedding representation; Step 5: Perform interactive iteration on the updated role embedding representation and the updated word embedding representation to obtain the current iteratively interacted role embedding representation and the current iteratively interacted word embedding representation respectively; Step 6: Repeat Step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation; Step 7: Input the final role embedding representation into a multi-layer perceptron for fusion and update, and use a span predictor for prediction to obtain a prediction result; Construct a cross-entropy loss based on the prediction result, optimize the prediction model using the cross-loss to obtain an optimized prediction model, and use the optimized prediction model to obtain the final prediction result.
[0005] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention proposes a dual self-attention driven embedding learning framework: by using a bidirectional query mechanism of role-word and word-role, it deeply explores the interaction semantics between roles at the pattern level and texts (words) at the instance level; through the dual self-attention mechanism, the model can more accurately capture the semantic information of roles from texts (words) and the semantic information of the roles played by words in events from event patterns (roles), and use them to update role embeddings and text (word) embeddings respectively; 2. The present invention designs a query vector, a key vector, and their interactive iterative update strategy: for the queries of the two self-attention mechanisms, combining the key features of roles and the context semantic information of words in texts, it designs an interactive iterative update strategy between the query vector and the key vector, and between role embeddings and word embeddings, effectively implementing the dual self-attention driven embedding learning method; 3. The present invention designs a model and a system that integrate the embedding learning framework and the embedding learning method to train and learn effective role and word representations.
[0006] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. Description of the Drawings
[0007] Figure 1 It is a flowchart of the steps of the method for learning role and word representations driven by dual self-attention proposed by the present invention.
[0008] Figure 2 It is a flow framework diagram of the method for learning role and word representations driven by dual self-attention proposed by the present invention.
[0009] Figure 3 It is an example diagram of the method for learning role and word representations driven by dual self-attention proposed by the present invention. Detailed Embodiments
[0010] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0011] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will become clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed as some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0012] Please refer to Figure 1 , an embodiment of the present invention proposes a method for role and word representation learning driven by dual self-attention, and the method includes the following steps: Step 1: Construct a feature fusion module based on a linear transformation mechanism and a gated neural network, and respectively construct a role-word self-attention module and a word-role self-attention module based on a bidirectional self-attention mechanism. The feature fusion module, the role-word self-attention module, the word-role self-attention module, a multi-layer perceptron, and a span predictor constitute a prediction model.
[0013] Step 2: Input the text and given a passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings.
[0014] Please refer to Figure 2 , in Step 2, input the text and given a passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings, which specifically include the following steps: Input the text and given a passage, perform word segmentation analysis and syntactic analysis on the passage in turn using a syntactic analysis tool, and perform word segmentation processing and random initialization operation on the passage in turn using a pre-trained language model to obtain the embedding of the trigger word, the embedding of the word semantics, the part-of-speech embedding, the dependency relationship type embedding, and the event type embedding respectively. The relational expressions existing in the corresponding processes are as follows: ; Among them, represents the embedding representation of the trigger word, represents the pre-trained language model, represents the trigger word in the passage, represents the th word semantic embedding, represents the part-of-speech embedding, Denote the embedding of dependency relation type, Denote the embedding of event type, Denote the embedding of role; Perform a marking insertion operation on the trigger word in the passage to obtain the passage after insertion marking. Input the passage after insertion marking into the pre-trained language model for encoding through the encoder to obtain the output passage embedding. Decode the output passage embedding through the decoder to obtain the passage embedding integrating the context semantics. The relational expressions existing in the corresponding process are as follows: ; Among them, Denote the passage embedding representation output by the pre-trained language model encoder representation, Denote the passage after passing through the pre-trained language model encoder and the decoder and the generated embedding representation; Given the event type hint for the passage, input the output passage embedding and the event type hint into the decoder for decoding to obtain the hint embedding containing the context semantics of the passage and the semantic of the event pattern. The relational expressions existing in the corresponding process are as follows: ; Among them, Denote the hint embedding containing the context semantics of the passage and the semantic of the event pattern, Denote the event type hint text.
[0015] Step 3: Use the feature fusion module to capture the span correlation between the pattern semantics of the role and the context semantics of the words in different semantic embeddings to obtain the initial embedding representation of the role and the initial embedding representation of the words.
[0016] In Step 3, use the feature fusion module to capture the span correlation between the pattern semantics of the role and the context semantics of the words in different semantic embeddings to obtain the initial embedding representation of the role and the initial embedding representation of the words, which specifically includes the following steps: Input the embedding of the role, the embedding of the event type, and the embedding of the trigger word into the feature fusion module and perform vector fusion processing in combination with the span predictor to respectively obtain the query vector of the role in the span start position predictor and the query vector of the role in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, Denote the query vector of the role in the span start position predictor, Denote the query vector of the role in the span end position predictor, Indicates that after being processed by the feature fusion module, represents the learnable weight matrix of the pattern semantics and context semantics in the span start position predictor, represents the learnable weight matrix of the pattern semantics and context semantics in the span end position predictor; Input the embedding of word semantics, part-of-speech embedding, and dependency relation type embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span start position predictor. Perform linear transformation processing on the key vector of the word in the span start position predictor to obtain the value vector of the word in the span start position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span start position predictor, represents the value vector of the word in the span start position predictor, represents being processed by linear transformation; Input the embedding of word semantics, part-of-speech embedding, and dependency relation type embedding into the feature fusion module again, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span end position predictor. Perform linear transformation processing on the key vector of the word in the span end position to obtain the value vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span end position predictor, represents the value vector of the word in the span end position predictor; Input the embedding of word semantics, part-of-speech embedding, and dependency relation type embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to respectively obtain the query vector of the word in the span start position predictor and the query vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the query vector of the word in the span start position predictor, represents the query vector of the word in the span end position predictor; Input the embeddings of the role, the event type embedding, and the trigger word embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively. Perform linear transformation processing on the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively to obtain the value vector of the role in the span start position predictor and the value vector of the role in the span end position predictor respectively. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the role in the span start position predictor, represents the value vector of the role in the span start position predictor, represents the key vector of the role in the span end position predictor, represents the value vector of the role in the span end position predictor.
[0017] Step 4: Based on the role-word self-attention module, use the dual self-attention mechanism to calculate the attention weights of the role-word, and update the initial embedding representation of the role to obtain the updated role embedding representation; Based on the word-role self-attention module, use the dual self-attention mechanism to calculate the attention weights of the word-role, and update the initial embedding representation of the word to obtain the updated word embedding representation.
[0018] Please refer to Figure 3 , in Step 4, use the dual self-attention mechanism to calculate the attention weights of the role-word, and update the initial embedding representation of the role to obtain the updated role embedding representation, which specifically includes the following steps: Perform multi-head self-attention operation on the query vector of the role in the span start position predictor in combination with the key vector and value vector of the word in the span start position predictor to obtain the attention weight matrix of the role in the span start position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the attention weight matrix of the role in the span start position predictor, represents the multi-head operation, represents being subjected to normalization processing, represents the passage all words in the matrix formed by the combination of the key vectors of the span start position predictor, represents the transpose operation of the matrix, represents the passage All the words in A matrix of value vector combinations in the span start position predictor, Represents the key vector of the word in the span start position predictor of dimension; Perform a multi-head self-attention operation on the query vector of the role in the span end position predictor, combining the key vector of the word in the span end position predictor and the value vector of the word in the span end position predictor, to obtain the attention weight matrix of the role in the span end position predictor. The corresponding relationship in the process is as follows: ; Where, Represents the attention weight matrix of the role in the span end position predictor, Represents the passage All the words in A matrix of key vector combinations of all the words in the span end position predictor, Represents the passage All the words in A matrix of value vector combinations of all the words in the span end position predictor; Perform a multiplication operation and a non-linear activation process on the attention weight matrix of the role in the span start position predictor and the value vector of the word in the span start position predictor in sequence, to obtain the updated query vector of the role in the span start position predictor. The corresponding relationship in the process is as follows: ; Where, Represents the updated query vector of the role in the span start position predictor, Represents after non-linear activation processing; Perform a multiplication operation and a non-linear activation process on the attention weight matrix of the role in the span end position predictor and the value vector of the word in the span end position predictor in sequence, to obtain the updated query vector of the role in the span end position predictor. The corresponding relationship in the process is as follows: ; Where, Represents the updated query vector of the role in the span end position predictor.
[0019] Use the dual self-attention mechanism to calculate the word-role attention weight and update the initial embedding representation of the word to obtain the updated word embedding representation. The specific steps are as follows: The query vector of the word in the span start position predictor is combined with the key vector of the role in the span start position predictor and the value vector of the role in the span start position predictor, and a multi-head self-attention operation is performed to obtain the attention weight matrix of the word in the span start position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the attention weight matrix of the word in the span start position predictor, represents the role The key vector of in the span start position predictor The dimension of, represents the specified event type All roles on The matrix formed by combining the key vectors of in the span start position predictor, represents the specified event type All roles on The matrix formed by combining the value vectors of in the span start position predictor; The query vector of the word in the span end position predictor is combined with the key vector of the role in the span end position predictor and the value vector of the role in the span end position predictor, and a multi-head self-attention operation is performed to obtain the attention weight matrix of the word in the span end position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the attention weight matrix of the word in the span end position predictor, represents the specified event type All roles on The matrix formed by combining the key vectors of in the span end position predictor, represents the specified event type All roles on The matrix formed by combining the value vectors of in the span end position predictor; The attention weight matrix of the word in the span start position predictor and the value vector of the role in the span start position predictor are successively multiplied and non-linearly activated to obtain the updated query vector of the word in the span start position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the updated query vector of the word in the span start position predictor; Multiply the attention weight matrix of the word in the span end position predictor and the value vector of the role in the span end position predictor in sequence, and perform non-linear activation processing to obtain the updated query vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the updated query vector of the word in the span end position predictor.
[0020] Step 5: Perform interactive iteration on the updated role embedding representation and the updated word embedding representation to obtain the current iteratively interactive role embedding representation and the current iteratively interactive word embedding representation respectively.
[0021] In Step 5, perform interactive iteration on the updated role embedding representation and the updated word embedding representation to obtain the current iteratively interactive role embedding representation and the current iteratively interactive word embedding representation respectively. The specific steps are as follows: Perform linear transformation processing on the updated query vector of the role in the span start position predictor and the updated query vector of the word in the span start position predictor respectively to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction; Perform linear transformation processing on the updated query vector of the role in the span end position predictor and the updated query vector of the word in the span end position predictor respectively to obtain the query vector of the role in the span end position predictor after iterative interaction and the query vector of the word in the span end position predictor after iterative interaction; Use the query vector of the role in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span start position predictor to obtain the key vector of the role in the span start position predictor after interactive iteration; Use the query vector of the role in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span end position predictor to obtain the key vector of the role in the span end position predictor after interactive iteration; Use the query vector of the word in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span start position predictor to obtain the key vector of the word in the span start position predictor after interactive iteration; Use the query vector of the word in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span end position predictor to obtain the key vector of the word in the span end position predictor after interactive iteration.
[0022] The query vector of the updated role in the span start position predictor and the query vector of the updated word in the span start position predictor are respectively subjected to linear transformation processing to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction. The relational expressions for the corresponding processes are as follows: ; Among them, represents the query vector of the role in the span start position predictor after the -th iteration, represents the query vector of the word in the span start position predictor after the -th iteration, represents the query vector of the role in the span start position predictor after the -th iteration, represents the query vector of the word in the span start position predictor after the -th iteration; In the step of using the query vector of the role in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span start position predictor to obtain the key vector of the role in the span start position predictor after interactive iteration, the relational expressions for the corresponding processes are as follows: ; Among them, represents the key vector of the word in the span start position predictor after the -th iteration; In the step of using the query vector of the word in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span start position predictor to obtain the key vector of the word in the span start position predictor after interactive iteration, the relational expressions for the corresponding processes are as follows: ; Among them, represents the key vector of the word in the span end position predictor after the -th iteration.
[0023] Step 6: Repeat Step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation.
[0024] In Step 6, repeating Step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation specifically includes the following steps: Use the key vector of the role in the span start position predictor after interactive iteration as the input in the next round, and iteratively interact with the query vector of the role in the span start position predictor after interactive iteration to obtain the query vector of the role in the span start position predictor after the next round of interactive iteration; Use the key vector of the role in the span end position predictor after interactive iteration as the input in the next round, and iteratively interact with the query vector of the role in the span end position predictor after interactive iteration to obtain the query vector of the role in the span end position predictor after the next round of interactive iteration; Use the key vector of the word in the span start position predictor after interactive iteration as the input in the next round, and iteratively interact with the query vector of the word in the span start position predictor after interactive iteration to obtain the query vector of the word in the span start position predictor after the next round of interactive iteration; Use the key vector of the word in the span end position predictor after interactive iteration as the input in the next round, and iteratively interact with the query vector of the word in the span end position predictor after interactive iteration to obtain the query vector of the word in the span end position predictor after the next round of interactive iteration.
[0025] Step 7: Input the final role embedding representation into a multi-layer perceptron for fusion and update, and use the span predictor to make a prediction to obtain a prediction result; Construct a cross-entropy loss based on the prediction result, use the cross loss to optimize the prediction model to obtain an optimized prediction model, and use the optimized prediction model to obtain the final prediction result.
[0026] In step 7, input the final role embedding representation into a multi-layer perceptron for fusion and update, and use the span predictor to make a prediction to obtain a prediction result, which specifically includes the following steps: Extract the role embedding from the hint embedding containing the context semantics of the passage and the semantics of the event pattern to obtain the embedding representation of the role extracted from the hint embedding; Input the embedding representation of the role extracted from the hint embedding and the query vector of the role in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the role in the span start position predictor. The corresponding relationship in the process is as follows: ; Among them, represents the embedding of the role in the span start position predictor represents the multi-layer perceptron, represents the hint embedding containing the context semantics of the passage and the semantics of the event pattern in the The embedded representation of a role indicating after rounds of iteration, the query vector of the role in the span start position predictor ; Input the embedded representation of the role extracted from the prompt embedding and the query vector of the role in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedded representation of the role in the span end position predictor; Input the passage embedding incorporating context semantics and the query vector of the word in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated passage embedding in the span start position predictor. The corresponding relationship in the process is as follows: ; wherein represents the embedding of the passage in the span start position predictor ; indicating after rounds of iteration, the query vector of the passage in the span start position predictor ; Input the passage embedding incorporating context semantics and the query vector of the word in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated passage embedding in the span end position predictor; For the updated embedded representation of the role in the span start position predictor, the updated embedded representation of the role in the span end position predictor, the updated passage embedding in the span start position predictor, and the updated passage embedding in the span end position predictor, respectively construct matrices to obtain the learnable weight matrix of the role in the span start position predictor, the learnable weight matrix of the role in the span end position predictor, the learnable weight matrix of the passage in the span start position predictor, and the learnable weight matrix of the passage in the span end position predictor; Use the learnable weight matrix of the role in the span start position predictor to perform a multiplication calculation on the updated embedded representation of the role in the span start position predictor. The corresponding relationship in the process is as follows: ; wherein represents the embedded representation of the role in the predicted span start position predictor represents the learnable weight matrix of the role in the span start position predictor in the prediction stage; Using the learnable weight matrix of the constructed role in the span end position predictor, multiply the embeddings of the role in the updated span end position predictor to obtain the role embeddings in the predicted span end position predictor; Using the learnable weight matrix of the constructed passage in the span start position predictor, multiply the embeddings of the passage in the updated span start position predictor to obtain the passage embeddings in the predicted span start position predictor; Using the learnable weight matrix of the constructed passage in the span end position predictor, multiply the embeddings of the passage in the updated span end position predictor to obtain the passage embeddings in the predicted span end position predictor; Using the role embeddings in the predicted span start position predictor, perform position prediction on the passage embeddings in the predicted span start position predictor to obtain the distribution of the predicted span start for the role on the passage. The corresponding relationship in the process is as follows: ; Among them, represents the distribution of the predicted span start for the role on the passage ; Using the role, embeddings in the predicted span end position predictor, perform position prediction on the passage embeddings in the predicted span end position predictor to obtain the distribution of the predicted span end for the role on the passage. The corresponding relationship in the process is as follows: ; Among them, represents the distribution of the predicted span end for the role on the passage ;
[0027] Based on the prediction results, construct the cross-entropy loss. The corresponding relationship in the process is as follows: ; Among them, represents the probability distribution in the span start position predictor of the role , represents the probability distribution in the span end position predictor of the role , represents the cross-entropy loss, represents the logarithmic function, represents the th role of the given event type, represents the total number of passages, represents the number of roles of the given event type, represents the number of documents.
[0028] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0029] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0030] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for learning role and word representation based on dual self-attention, characterized in that: The method comprises the following steps: Step 1: construct a feature fusion module based on the linear transformation mechanism and the gated neural network, and construct a role-word self-attention module and a word-role self-attention module based on the bidirectional self-attention mechanism. The feature fusion module, the role-word self-attention module, the word-role self-attention module, the multi-layer perceptron and the span predictor constitute a prediction model; Step 2: Input text and give a chapter, perform syntactic and semantic analysis on the chapter, extract multi-level features with the pre-trained language model, and obtain different semantic embeddings; Step 3: Use the feature fusion module to capture the span association of the role's pattern semantics and the word's context semantics for different semantic embeddings, and obtain the initial embedding representation of the role and the initial embedding representation of the word; Step 4: Based on the role-word self-attention module, the dual-element self-attention mechanism is used to calculate the role-word attention weight, and the initial embedding representation of the role is updated to obtain the updated role embedding representation; Based on the word-role self-attention module, the attention weight of the word-role is calculated using the binary self-attention mechanism, and the initial embedding representation of the word is updated to obtain the updated word embedding representation; Step 5: interactively iterate the updated role embedding representation and the updated word embedding representation to obtain the role embedding representation after the current interactive iteration and the word embedding representation after the current interactive iteration respectively; Step 6: Repeat step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation; Step 7: Input the final role embedding representation into the multi-layer perceptron for fusion update, and use the span predictor to make predictions to obtain the prediction results; Based on the prediction results, a cross entropy loss is constructed, and the prediction model is optimized using the cross entropy loss to obtain an optimized prediction model, and the final prediction result is obtained using the optimized prediction model.
2. The method for learning role and word representation based on dual self-attention drive according to claim 1, characterized in that: In step 2, a text is input and a chapter is given, syntactic and semantic analysis is performed on the chapter, and multi-level features are extracted in combination with a pre-trained language model to obtain different semantic embeddings, which specifically includes the following steps: Input text and give a chapter, use the syntactic analysis tool to perform word segmentation analysis and syntactic analysis on the chapter in sequence, and use the pre-trained language model to perform word segmentation and random initialization operations in sequence, and obtain the embedding of trigger words, word semantics, part of speech embedding, dependency type embedding and event type embedding respectively. The relationship between the corresponding processes is as follows: ; in, represents the embedding representation of the trigger word, represents the pre-trained language model, Indicates the trigger words in the passage. Indicates Words The semantic embedding of represents part-of-speech embedding, represents a random initialization operation, represents dependency type embedding, Indicates event type embedding, Representing roles Embedding The trigger words in the article are marked and inserted to obtain the marked article. The marked article is input into the pre-trained language model and encoded by the encoder to obtain the output article embedding. The output article embedding is decoded by the decoder to obtain the article embedding that incorporates the contextual semantics. The corresponding process has the following relationship: ; in, Represents a pre-trained language model encoder Output chapter Embedding means, Indicates chapter Pre-trained language model encoder With decoder The generated embedding representation; Given an event type prompt for the passage, the output passage embedding and event type prompt are input into the decoder for decoding, and the prompt embedding containing the passage context semantics and event pattern semantics is obtained. The corresponding process has the following relationship: ; in, Represents the hint embedding that contains the text context semantics and event pattern semantics, Indicates the event type hint text.
3. The role and word representation learning method based on dual self-attention driving according to claim 2 is characterized in that: In step 3, the feature fusion module is used to capture the span association of the role's pattern semantics and the word's context semantics for different semantic embeddings, and the initial embedding representation of the role and the initial embedding representation of the word are obtained, which specifically includes the following steps: The role embedding, event type embedding and trigger word embedding are input into the feature fusion module, and the vector fusion processing is performed in combination with the span predictor to obtain the query vector of the role in the span start position predictor and the query vector of the role in the span end position predictor. The corresponding process has the following relationship: ; in, represents the query vector of the character in the span start position predictor, represents the query vector of the character in the span end position predictor, It means that it has been processed by the feature fusion module. A learnable weight matrix representing the pattern semantics and context semantics in the span start position predictor, A learnable weight matrix representing the pattern semantics and context semantics in the span end position predictor; The word semantic embedding, part-of-speech embedding, and dependency type embedding are input into the feature fusion module, and the vector fusion processing is performed in combination with the span predictor to obtain the key vector of the word in the span start position predictor. The key vector of the word in the span start position predictor is linearly transformed to obtain the value vector of the word in the span start position predictor. The corresponding process has the following relationship: ; in, represents the key vector of the word in the span start position predictor, represents the value vector of the word in the span start position predictor, Indicates that it has been processed by linear transformation; The word semantic embedding, part-of-speech embedding, and dependency type embedding are input into the feature fusion module again, and vector fusion processing is performed in combination with the span predictor to obtain the key vector of the word in the span end position predictor. The key vector of the word at the span end position is linearly transformed to obtain the value vector of the word in the span end position predictor. The corresponding process has the following relationship: ; in, represents the key vector of the word in the span end position predictor, A vector representing the value of a word in the span end position predictor; The word semantic embedding, part-of-speech embedding, and dependency type embedding are input into the feature fusion module, and vector fusion processing is performed in combination with the span predictor to obtain the query vector of the word in the span start position predictor and the query vector of the word in the span end position predictor. The corresponding process has the following relationship: ; in, represents the query vector of the word in the span start position predictor, The query vector representing the word in the span end position predictor; The role embedding, event type embedding and trigger word embedding are input into the feature fusion module, and the vector fusion processing is performed in combination with the span predictor to obtain the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively. The key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor are linearly transformed to obtain the value vector of the role in the span start position predictor and the value vector of the role in the span end position predictor respectively. The relationship between the corresponding processes is as follows: ; in, represents the key vector of the role in the span start position predictor, represents the value vector of the character in the span start position predictor, represents the key vector of the role in the span end position predictor, A vector of values representing the role's span end position predictor.
4. The role and word representation learning method based on dual self-attention driving according to claim 3 is characterized in that: In step 4, the attention weight of the role-word is calculated using the binary self-attention mechanism, and the initial embedding representation of the role is updated to obtain an updated role embedding representation, which specifically includes the following steps: For the query vector of the character in the span start position predictor, combined with the key vector of the word in the span start position predictor and the value vector of the word in the span start position predictor, a multi-head self-attention operation is performed to obtain the attention weight matrix of the character in the span start position predictor. The corresponding process has the following relationship: ; in, represents the attention weight matrix of the character in the span start position predictor, Indicates long position operation. It means that after normalization, Indicates chapter All words in a matrix of key vector combinations in the span start position predictor, represents the transpose operation of the matrix, Indicates chapter All words in A matrix of combinations of value vectors in the span start position predictor, The key vector representing the word in the span start position predictor Dimensions; For the query vector of the character in the span end position predictor, combined with the key vector of the word in the span end position predictor and the value vector of the word in the span end position predictor, a multi-head self-attention operation is performed to obtain the attention weight matrix of the character in the span end position predictor. The corresponding process has the following relationship: ; in, represents the attention weight matrix of the character in the span end position predictor, Indicates chapter All words in a matrix of key vector combinations in the span end position predictor, Indicates chapter All words in A matrix of combinations of value vectors in the span end position predictor; The attention weight matrix of the character in the span start position predictor and the value vector of the word in the span start position predictor are multiplied and nonlinearly activated in sequence to obtain the updated query vector of the character in the span start position predictor. The corresponding process has the following relationship: ; in, represents the query vector of the updated character in the span start position predictor, Indicates that it has been processed by nonlinear activation; The attention weight matrix of the character in the span end position predictor and the value vector of the word in the span end position predictor are multiplied and activated nonlinearly in sequence to obtain the updated query vector of the character in the span end position predictor. The corresponding process has the following relationship: ; in, Query vector representing the updated character in the span end position predictor.
5. The method for learning role and word representation based on dual self-attention drive according to claim 4, characterized in that: The binary self-attention mechanism is used to calculate the word-role attention weight, and the initial embedding representation of the word is updated to obtain the updated word embedding representation, which specifically includes the following steps: The query vector of the word in the span start position predictor is combined with the key vector of the role in the span start position predictor and the value vector of the role in the span start position predictor to perform a multi-head self-attention operation to obtain the attention weight matrix of the word in the span start position predictor. The corresponding process has the following relationship: ; in, represents the attention weight matrix of the word in the span start position predictor, Representing roles Key vector in span start predictor The dimension of Indicates the specified event type All roles a matrix of key vector combinations in the span start position predictor, Indicates the specified event type All roles A matrix of combinations of value vectors in the span start position predictor; For the query vector of the word in the span end position predictor, combined with the key vector of the role in the span end position predictor and the value vector of the role in the span end position predictor, a multi-head self-attention operation is performed to obtain the attention weight matrix of the word in the span end position predictor. The corresponding process has the following relationship: ; in, represents the attention weight matrix of the word in the span end position predictor, Indicates the specified event type All roles a matrix of key vector combinations in the span end position predictor, Indicates the specified event type All roles A matrix of combinations of value vectors in the span end position predictor; The attention weight matrix of the word in the span start position predictor and the value vector of the role in the span start position predictor are multiplied and nonlinearly activated in sequence to obtain the updated query vector of the word in the span start position predictor. The corresponding process has the following relationship: ; in, The query vector representing the updated word in the span start position predictor; The attention weight matrix of the word in the span end position predictor and the value vector of the role in the span end position predictor are multiplied and nonlinearly activated in sequence to obtain the updated query vector of the word in the span end position predictor. The corresponding process has the following relationship: ; in, The query vector representing the updated term in the span end position predictor.
6. The method for learning role and word representation based on dual self-attention drive according to claim 5, characterized in that: In step 5, interactive iteration is performed on the updated role embedding representation and the updated word embedding representation to obtain the role embedding representation after the current interactive iteration and the word embedding representation after the current interactive iteration, respectively, which specifically includes the following steps: Performing linear transformation processing on the updated query vector of the role in the span start position predictor and the updated query vector of the word in the span start position predictor, respectively, to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction; Performing linear transformation processing on the updated query vector of the role in the span end position predictor and the updated query vector of the word in the span end position predictor, respectively, to obtain the query vector of the role in the span end position predictor after iterative interaction and the query vector of the word in the span end position predictor after iterative interaction; Using the query vector of the character in the span start position predictor after the iterative interaction, interactively iterate the key vector of the character in the span start position predictor to obtain the key vector of the character in the span start position predictor after the interactive iteration; Using the query vector of the character in the span end position predictor after the iterative interaction, interactively iterate the key vector of the character in the span end position predictor to obtain the key vector of the character in the span end position predictor after the interactive iteration; Using the query vector of the word in the span start position predictor after iterative interaction, interactively iterate the key vector of the word in the span start position predictor to obtain the key vector of the word in the span start position predictor after interactive iteration; The query vector of the word in the span end position predictor after the iterative interaction is used to interactively iterate the key vector of the word in the span end position predictor to obtain the key vector of the word in the span end position predictor after the interactive iteration.
7. The method for learning role and word representation based on dual self-attention drive according to claim 6, characterized in that: The updated query vector of the role in the span start position predictor and the updated query vector of the word in the span start position predictor are linearly transformed to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction. The corresponding process has the following relationship: ; in, Indicates that after Role in span start position predictor after round iteration The query vector is Indicates that after After rounds of iterations, the word in the span start position predictor The query vector is Indicates that after Role in span start position predictor after round iteration The query vector is Indicates that after After rounds of iterations, the word in the span start position predictor The query vector of In the step of interactively iterating the key vector of the character in the span start position predictor by using the query vector of the character in the span start position predictor after iterative interaction to obtain the key vector of the character in the span start position predictor after interactive iteration, the relationship between the corresponding process is as follows: ; in, Indicates that after The key vector of the word in the span start position predictor after round iteration; In the step of interactively iterating the key vector of the word in the span start position predictor by using the query vector of the word in the span start position predictor after iterative interaction to obtain the key vector of the word in the span start position predictor after interactive iteration, the relationship between the corresponding process is as follows: ; in, Indicates that after The key vector of the word in the span end position predictor after the round iteration.
8. The method for learning role and word representation based on dual-element self-attention drive according to claim 7, characterized in that: In step 6, step 5 is repeated in an iterative manner to obtain the final role embedding representation and the final word embedding representation, which specifically includes the following steps: Using the key vector of the character in the span start position predictor after the interactive iteration as input in the next round, interactively iterate the query vector of the character in the span start position predictor after the interactive iteration in an iterative manner to obtain the query vector of the character in the span start position predictor after the next round of interactive iteration; Using the key vector of the character in the span end position predictor after the interactive iteration as input in the next round, interactively iterate the query vector of the character in the span end position predictor after the interactive iteration in an iterative manner to obtain the query vector of the character in the span end position predictor after the next round of interactive iteration; Using the key vector of the word in the span start position predictor after the interactive iteration as input in the next round, interactively iterate the query vector of the word in the span start position predictor after the interactive iteration in an iterative manner to obtain the query vector of the word in the span start position predictor after the next round of interactive iteration; The key vector of the word in the span end position predictor after the interactive iteration is used as input in the next round, and the query vector of the word in the span end position predictor after the interactive iteration is interactively iterated in an iterative manner to obtain the query vector of the word in the span end position predictor after the next round of interactive iteration.
9. The method for learning role and word representation based on dual self-attention drive according to claim 8, characterized in that: In step 7, the final role embedding representation is input into the multi-layer perceptron for fusion update, and the span predictor is used for prediction to obtain the prediction result, which specifically includes the following steps: Extract role embedding from the prompt embedding that contains the text context semantics and event pattern semantics, and obtain the embedded representation of the role extracted from the prompt embedding; The embedding representation of the character extracted from the prompt embedding and the query vector of the character in the span start position predictor after the next round of interaction iteration are input into the multi-layer perceptron for fusion update to obtain the embedding of the character in the updated span start position predictor. The corresponding process has the following relationship: ; in, Represents the role in the span start position predictor The embedding represents a multi-layer perceptron, Representing hint embeddings that imply discourse context and event pattern semantics Middle The embedding representation of the role, Indicates passing Role in span start position predictor after round iteration The query vector of The embedded representation of the character extracted from the prompt embedding and the query vector of the character in the span end position predictor after the next round of interaction iteration are input into the multi-layer perceptron for fusion update to obtain the updated embedding of the character in the span end position predictor; The chapter embedding that incorporates contextual semantics and the query vector of the word in the span start position predictor after the next round of interactive iteration are input into the multi-layer perceptron for fusion update to obtain the embedding of the chapter in the updated span start position predictor. The corresponding process has the following relationship: ; in, Represents the span start position predictor in the chapter Embedded, Indicates passing After round iteration, in the span start position predictor The query vector of The chapter embedding that incorporates contextual semantics and the query vector of the word in the span end position predictor after the next round of interactive iteration are input into the multi-layer perceptron for fusion update to obtain the embedding of the chapter in the updated span end position predictor; Construct matrices for the embedding of the role in the updated span start position predictor, the embedding of the role in the updated span end position predictor, the embedding of the chapter in the updated span start position predictor, and the embedding of the chapter in the updated span end position predictor, respectively, to obtain a learnable weight matrix of the constructed role in the span start position predictor, a learnable weight matrix of the constructed role in the span end position predictor, a learnable weight matrix of the constructed chapter in the span start position predictor, and a learnable weight matrix of the constructed chapter in the span end position predictor; Using the constructed learnable weight matrix of the role in the span start position predictor, the embedding of the role in the updated span start position predictor is multiplied to obtain the role embedding in the predicted span start position predictor. The corresponding process has the following relationship: ; in, represents the role embedding in the prediction span start position predictor, Represents the learnable weight matrix of the span start position predictor for the role in the prediction phase; Using the constructed learnable weight matrix of the role in the span end position predictor, multiply the embedding of the role in the updated span end position predictor to obtain the embedding of the role in the predicted span end position predictor; Using the constructed learnable weight matrix of the chapter in the span start position predictor, multiply the embedding of the chapter in the updated span start position predictor to obtain the embedding of the chapter in the predicted span start position predictor; Using the constructed learnable weight matrix of the chapter in the span end position predictor, multiply the embedding of the chapter in the updated span end position predictor to obtain the embedding of the chapter in the predicted span end position predictor; Using the role embedding in the predicted span start position predictor, the chapter embedding in the predicted span start position predictor is predicted to obtain the distribution of the span start predicted for the role in the chapter. The corresponding process has the following relationship: ; in, Indicates the role In the chapter The distribution of the span started on the forecast; Using the role embedding in the predicted span end position predictor, the chapter embedding in the predicted span end position predictor is predicted to obtain the distribution of the span end predicted for the role in the chapter. The corresponding process has the following relationship: ; in, Indicates the role In the chapter The distribution at the end of the upper forecast span.
10. The method for learning role and word representation based on dual self-attention drive according to claim 9, characterized in that: Based on the prediction results, the cross entropy loss is constructed, and the relationship between the corresponding process is as follows: ; in, Representing roles The probability distribution in the span start position predictor, Representing roles The probability distribution in the span end position predictor, represents the cross entropy loss, represents the logarithmic function, Indicates the first roles, Indicates the number of all chapters, represents the number of roles for a given event type, Indicates the number of documents.
Citation Information
Patent Citations
Multi-task chapter-level event extraction method based on multi-headed self-attention mechanism
CN113761936A
Method for identifying chapter-level event roles based on knowledge graph information guidance
CN114880434A
Semantic role labeling method based on part-of-speech and attention fusion
CN117933266A
Relying on discourse analysis to answer complex questions by neural machine reading comprehension
US20220138432A1
Method and apparatus for recognizing role in text, and readable medium and electronic device
WO2022166613A1