A Role and Word Representation Learning Method Driven by Dual Self-Attention

Through the dual self-attention-driven role and word representation learning method, the problem that the semantic relationship between roles and words in the existing technology is not fully explored, and a more accurate information extraction effect is achieved.

CN120067306BActive Publication Date: 2025-07-08JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510562662.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing role-based span selection strategy fails to fully explore the relationship between the semantics of each role in the event type and the semantics of each word that act as arguments in the chapter, resulting in inaccurate information extraction.

Method used

A dual self-attention-driven character and word representation learning method is adopted. By constructing a role-word self-attention module and a word-role self-attention module, combining a multi-layer perceptron and span predictor, feature fusion and interactive iteration are performed to update the role and word embedding representation.

Benefits of technology

Deeply exploring the interactive semantics between characters and words, improving the ability to capture role semantic information, and enhancing the accuracy and efficiency of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067306B_ABST
    Figure CN120067306B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for learning character and word representations driven by dual self-attention, which includes: inputting text and a given passage, performing syntactic and semantic analysis on the passage, and extracting multi-level features in combination with a pre-trained language model; using a feature fusion module to capture the span correlation of the pattern semantics of characters and the context semantics of words for different semantic embeddings; using a bidirectional self-attention mechanism to calculate the attention weights of character-word and word-character, and updating the initial embedding representations of characters and the initial embeddings of words; performing interactive iteration on the updated character embedding representations and the updated word embedding representations; inputting the final character embedding representations into a multi-layer perceptron for fusion and update, and using a span predictor for prediction to obtain a prediction result. The present invention designs a model and a system that integrate an embedding learning framework and an embedding learning method, and trains and learns effective character and word representations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information extraction, and particularly to a method for learning the representation of roles and words driven by dual self-attention. Background Art

[0002] Event argument extraction at the discourse level is a key task in information extraction, which extracts event-related arguments from a discourse and accurately identifies their roles. The existing role-based span selection strategy mainly captures role representations and discourse representations through pre-trained language models and prompts, so as to construct a span selector to predict spans for each role. However, the existing methods do not fully explore the association between the semantics of each role in the event type and the semantics of each word serving as an argument in the discourse. Summary of the Invention

[0003] In view of the above situation, the main purpose of the present invention is to propose a method for learning the representation of roles and words driven by dual self-attention to solve the above technical problems.

[0004] The present invention proposes a method for learning the representation of roles and words driven by dual self-attention, and the method includes the following steps:

[0005] Step 1: Construct a feature fusion module based on a linear transformation mechanism and a gated neural network, and respectively construct a role-word self-attention module and a word-role self-attention module based on a bidirectional self-attention mechanism. The feature fusion module, the role-word self-attention module, the word-role self-attention module, a multi-layer perceptron, and a span predictor constitute a prediction model;

[0006] Step 2: Input a text and a given discourse, perform syntactic and semantic analysis on the discourse, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings;

[0007] Step 3: Use the feature fusion module to capture the span association between the pattern semantics of roles and the context semantics of words for different semantic embeddings, and obtain the initial embedding representation of roles and the initial embedding representation of words;

[0008] Step 4: Based on the role-word self-attention module, use the bidirectional self-attention mechanism to calculate the attention weights of role-words, and update the initial embedding representation of roles to obtain an updated role embedding representation;

[0009] Based on the word-role self-attention module, use the bidirectional self-attention mechanism to calculate the attention weights of word-roles, and update the initial embedding representation of words to obtain an updated word embedding representation;

[0010] Step 5: Interactively iterate the updated role embedding representation and the updated word embedding representation to obtain the role embedding representation after the current interactive iteration and the word embedding representation after the current interactive iteration respectively;

[0011] Step 6: Repeat Step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation;

[0012] Step 7: Input the final role embedding representation into a multi-layer perceptron for fusion and update, and use a span predictor for prediction to obtain a prediction result;

[0013] Construct a cross-entropy loss based on the prediction result, optimize the prediction model using the cross loss to obtain an optimized prediction model, and use the optimized prediction model to obtain the final prediction result.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0015] 1. The present invention proposes a dual self-attention driven embedding learning framework: by using a bidirectional query mechanism of role-word and word-role, it deeply explores the interactive semantics between roles at the pattern level and texts (words) at the instance level; through the dual self-attention mechanism, the model can more accurately capture the semantic information of roles from texts (words) and the semantic information of words playing roles in events from event patterns (roles), and use them to update role embeddings and text (word) embeddings respectively;

[0016] 2. The present invention designs a query vector, a key vector and their interactive iterative update strategy: for the queries of the two self-attention mechanisms, combining the key features of roles and the context semantic information of words in texts, it designs an interactive iterative update strategy between the query vector and the key vector, and between role embeddings and word embeddings, effectively realizing the dual self-attention driven embedding learning method;

[0017] 3. The present invention designs a model and a system that integrate the embedding learning framework and the embedding learning method to train and learn effective role and word representations.

[0018] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0019] Figure 1 It is a flowchart of the steps of the method for learning role and word representations driven by dual self-attention proposed by the present invention.

[0020] Figure 2 It is a flow framework diagram of the method for learning role and word representations driven by dual self-attention proposed by the present invention.

[0021] Figure 3 This is an example diagram of the method for learning character and word representations driven by dual self-attention proposed by the present invention. Detailed implementation manners

[0022] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0023] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention. However, it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0024] Please refer to Figure 1 , the embodiments of the present invention propose a method for learning character and word representations driven by dual self-attention. The method includes the following steps:

[0025] Step 1: Construct a feature fusion module based on a linear transformation mechanism and a gated neural network, and respectively construct a character-word self-attention module and a word-character self-attention module based on a bidirectional self-attention mechanism. The feature fusion module, the character-word self-attention module, the word-character self-attention module, a multi-layer perceptron, and a span predictor constitute a prediction model.

[0026] Step 2: Input the text and given a passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings.

[0027] Please refer to Figure 2 , in Step 2, input the text and given a passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings, which specifically include the following steps:

[0028] Input the text and given a passage, perform word segmentation analysis and syntactic analysis on the passage in turn using a syntactic analysis tool, and perform word segmentation processing and random initialization operation on the passage in turn using a pre-trained language model to obtain the embedding of the trigger word, the embedding of the word semantics, the part-of-speech embedding, the dependency relationship type embedding, and the event type embedding respectively. The relational expressions existing in the corresponding processes are as follows:

[0029] ;

[0030] Among them, represents the embedding representation of the trigger word, denotes a pre-trained language model, denotes the trigger word in the passage, denotes the th word semantic embedding, denotes the part-of-speech embedding, denotes the random initialization operation, denotes the dependency relation type embedding, denotes the event type embedding, denotes the role embedding;

[0031] Perform a marker insertion operation on the trigger word in the passage to obtain the passage with inserted markers. Input the passage with inserted markers into the pre-trained language model for encoding through the encoder to obtain the output passage embedding. Decode the output passage embedding through the decoder to obtain the passage embedding incorporating context semantics. The corresponding relationship in the process is as follows:

[0032] ;

[0033] where, denotes the passage embedding representation output by the pre-trained language model encoder representation, denotes the passage after passing through the pre-trained language model encoder and the decoder and the generated embedding representation;

[0034] Given an event type hint for the passage, input the output passage embedding and the event type hint into the decoder for decoding to obtain a hint embedding containing passage context semantics and event pattern semantics. The corresponding relationship in the process is as follows:

[0035] ;

[0036] where, denotes the hint embedding containing passage context semantics and event pattern semantics, denotes the event type hint text.

[0037] Step 3: Use the feature fusion module to capture the span correlation between the pattern semantics of the role and the context semantics of the words in different semantic embeddings, and obtain the initial embedding representation of the role and the initial embedding representation of the words.

[0038] In Step 3, use the feature fusion module to capture the span correlation between the pattern semantics of the role and the context semantics of the words in different semantic embeddings, and obtain the initial embedding representation of the role and the initial embedding representation of the words. The specific steps are as follows:

[0039] Input the embeddings of the role, the event type embedding, and the trigger word embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to respectively obtain the query vector of the role in the span start position predictor and the query vector of the role in the span end position predictor. The relational expressions in the corresponding process are as follows:

[0040] ;

[0041] Among them, represents the query vector of the role in the span start position predictor, represents the query vector of the role in the span end position predictor, represents being processed by the feature fusion module, represents the learnable weight matrix of the pattern semantics and context semantics in the span start position predictor, represents the learnable weight matrix of the pattern semantics and context semantics in the span end position predictor;

[0042] Input the embedding of word semantics, part-of-speech embedding, and dependency relation type embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span start position predictor. Perform a linear transformation on the key vector of the word in the span start position predictor to obtain the value vector of the word in the span start position predictor. The relational expressions in the corresponding process are as follows:

[0043] ;

[0044] Among them, represents the key vector of the word in the span start position predictor, represents the value vector of the word in the span start position predictor, represents being processed by a linear transformation;

[0045] Input the embedding of word semantics, part-of-speech embedding, and dependency relation type embedding into the feature fusion module again, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span end position predictor. Perform a linear transformation on the key vector of the word in the span end position to obtain the value vector of the word in the span end position predictor. The relational expressions in the corresponding process are as follows:

[0046] ;

[0047] Among them, represents the key vector of the word in the span end position predictor, represents the value vector of the word in the span end position predictor;

[0048] Embed the word semantics embedding, part-of-speech embedding, and dependency relation type embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the query vector of the word in the span start position predictor and the query vector of the word in the span end position predictor respectively. The relational expressions for the corresponding processes are as follows:

[0049] ;

[0050] Among them, represents the query vector of the word in the span start position predictor, represents the query vector of the word in the span end position predictor;

[0051] Embed the role embedding, event type embedding, and trigger word embedding into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively. Perform linear transformation processing on the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively to obtain the value vector of the role in the span start position predictor and the value vector of the role in the span end position predictor respectively. The relational expressions for the corresponding processes are as follows:

[0052] ;

[0053] Among them, represents the key vector of the role in the span start position predictor, represents the value vector of the role in the span start position predictor, represents the key vector of the role in the span end position predictor, represents the value vector of the role in the span end position predictor.

[0054] Step 4: Based on the role-word self-attention module, use the dual self-attention mechanism to calculate the attention weights of the role-word, and update the initial embedding representation of the role to obtain the updated role embedding representation;

[0055] Based on the word-role self-attention module, use the dual self-attention mechanism to calculate the attention weights of the word-role, and update the initial embedding representation of the word to obtain the updated word embedding representation.

[0056] Please refer to Figure 3 , in Step 4, use the dual self-attention mechanism to calculate the attention weights of the role-word, and update the initial embedding representation of the role to obtain the updated role embedding representation, which specifically includes the following steps:

[0057] Perform a multi-head self-attention operation on the query vector of the role in the span start position predictor, combined with the key vector of the word in the span start position predictor and the value vector of the word in the span start position predictor, to obtain the attention weight matrix of the role in the span start position predictor. The corresponding relationship in the process is as follows:

[0058] ;

[0059] Among them, represents the attention weight matrix of the role in the span start position predictor, represents the multi-head operation, represents being processed by normalization, represents the passage all words in the matrix combined by the key vectors in the span start position predictor, represents the transpose operation of the matrix, represents the passage all words in the matrix combined by the value vectors in the span start position predictor, represents the key vector of the word in the span start position predictor of the dimension;

[0060] Perform a multi-head self-attention operation on the query vector of the role in the span end position predictor, combined with the key vector of the word in the span end position predictor and the value vector of the word in the span end position predictor, to obtain the attention weight matrix of the role in the span end position predictor. The corresponding relationship in the process is as follows:

[0061] ;

[0062] Among them, represents the attention weight matrix of the role in the span end position predictor, represents the passage all words in the matrix combined by the key vectors in the span end position predictor, represents the passage all words in the matrix combined by the value vectors in the span end position predictor;

[0063] Perform a multiplication operation and a non-linear activation process on the attention weight matrix of the role in the span start position predictor and the value vector of the word in the span start position predictor in sequence, to obtain the updated query vector of the role in the span start position predictor. The corresponding relationship in the process is as follows:

[0064] ;

[0065] Among them, represents the query vector of the updated role in the span start position predictor, indicating being processed by non-linear activation;

[0066] Perform multiplication operation and non-linear activation processing on the attention weight matrix of the role in the span end position predictor and the value vector of the word in the span end position predictor in sequence to obtain the query vector of the updated role in the span end position predictor. The relational expressions existing in the corresponding process are as follows:

[0067] ;

[0068] Among them, represents the query vector of the updated role in the span end position predictor.

[0069] Use the binary self-attention mechanism to calculate the word-role attention weights and update the initial embedding representation of the word to obtain the updated word embedding representation. The specific steps are as follows:

[0070] Perform multi-head self-attention operation on the query vector of the word in the span start position predictor, combined with the key vector of the role in the span start position predictor and the value vector of the role in the span start position predictor, to obtain the attention weight matrix of the word in the span start position predictor. The relational expressions existing in the corresponding process are as follows:

[0071] ;

[0072] Among them, represents the attention weight matrix of the word in the span start position predictor, represents the role in the key vector of the span start position predictor of the dimension, represents the specified event type on all roles in the key vector combination matrix of the span start position predictor, represents the specified event type on all roles in the value vector combination matrix of the span start position predictor;

[0073] Perform multi-head self-attention operation on the query vector of the word in the span end position predictor, combined with the key vector of the role in the span end position predictor and the value vector of the role in the span end position predictor, to obtain the attention weight matrix of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows:

[0074] ;

[0075] Among them, represents the attention weight matrix of the word in the span end position predictor, represents the specified event type for all roles a matrix of the combination of key vectors in the span end position predictor, represents the specified event type for all roles a matrix of the combination of value vectors in the span end position predictor;

[0076] Multiply the attention weight matrix of the word in the span start position predictor and the value vector of the role in the span start position predictor in sequence, and perform non-linear activation processing to obtain the updated query vector of the word in the span start position predictor. The relational expressions existing in the corresponding process are as follows:

[0077] ;

[0078] Among them, represents the updated query vector of the word in the span start position predictor;

[0079] Multiply the attention weight matrix of the word in the span end position predictor and the value vector of the role in the span end position predictor in sequence, and perform non-linear activation processing to obtain the updated query vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows:

[0080] ;

[0081] Among them, represents the updated query vector of the word in the span end position predictor.

[0082] Step 5: Interact and iterate the updated role embedding representation and the updated word embedding representation to obtain the current iterated role embedding representation and the current iterated word embedding representation respectively.

[0083] In Step 5, interact and iterate the updated role embedding representation and the updated word embedding representation to obtain the current iterated role embedding representation and the current iterated word embedding representation respectively, which specifically includes the following steps:

[0084] Perform linear transformation processing on the query vector of the updated role in the span start position predictor and the query vector of the updated word in the span start position predictor respectively, to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction;

[0085] Perform linear transformation processing on the query vector of the updated role in the span end position predictor and the query vector of the updated word in the span end position predictor respectively, to obtain the query vector of the role in the span end position predictor after iterative interaction and the query vector of the word in the span end position predictor after iterative interaction;

[0086] Use the query vector of the role in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span start position predictor, to obtain the key vector of the role in the span start position predictor after interactive iteration;

[0087] Use the query vector of the role in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span end position predictor, to obtain the key vector of the role in the span end position predictor after interactive iteration;

[0088] Use the query vector of the word in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span start position predictor, to obtain the key vector of the word in the span start position predictor after interactive iteration;

[0089] Use the query vector of the word in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span end position predictor, to obtain the key vector of the word in the span end position predictor after interactive iteration.

[0090] Perform linear transformation processing on the query vector of the updated role in the span start position predictor and the query vector of the updated word in the span start position predictor respectively, to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction. The relational expressions existing in the corresponding process are as follows:

[0091] ;

[0092] Among them, represents the query vector of the role in the span start position predictor after the -th round of iteration, represents the query vector of the word in the span start position predictor after the -th round of iteration, Denote the query vector of the role in the span start position predictor after the th iteration, Denote the query vector of the token in the span start position predictor after the th iteration;

[0093] In the step of iteratively interacting the query vector of the role in the span start position predictor after iterative interaction with the key vector of the role in the span start position predictor to obtain the key vector of the role in the span start position predictor after iterative interaction, the following relational expressions exist for the corresponding process:

[0094] ;

[0095] where, denotes the key vector of the token in the span start position predictor after the th iteration;

[0096] In the step of iteratively interacting the query vector of the token in the span start position predictor after iterative interaction with the key vector of the token in the span start position predictor to obtain the key vector of the token in the span start position predictor after iterative interaction, the following relational expressions exist for the corresponding process:

[0097] ;

[0098] where, denotes the key vector of the token in the span end position predictor after the th iteration.

[0099] Step 6: Repeat Step 5 iteratively to obtain the final role embedding representation and the final token embedding representation.

[0100] In Step 6, repeat Step 5 iteratively to obtain the final role embedding representation and the final token embedding representation, which specifically includes the following steps:

[0101] Use the key vector of the role in the span start position predictor after iterative interaction as the input in the next round, and iteratively interact the query vector of the role in the span start position predictor after iterative interaction to obtain the query vector of the role in the span start position predictor after the next round of iterative interaction;

[0102] Using the key vector of the role in the span end position predictor after interactive iteration as the input in the next round, interactively iterate the query vector of the role in the span end position predictor after interactive iteration in an iterative manner to obtain the query vector of the role in the span end position predictor after the next round of interactive iteration;

[0103] Using the key vector of the word in the span start position predictor after interactive iteration as the input in the next round, interactively iterate the query vector of the word in the span start position predictor after interactive iteration in an iterative manner to obtain the query vector of the word in the span start position predictor after the next round of interactive iteration;

[0104] Using the key vector of the word in the span end position predictor after interactive iteration as the input in the next round, interactively iterate the query vector of the word in the span end position predictor after interactive iteration in an iterative manner to obtain the query vector of the word in the span end position predictor after the next round of interactive iteration.

[0105] Step 7: Input the final role embedding representation into a multi-layer perceptron for fusion update, and use the span predictor to make a prediction to obtain a prediction result;

[0106] Construct a cross-entropy loss based on the prediction result, use the cross loss to optimize the prediction model to obtain an optimized prediction model, and use the optimized prediction model to obtain the final prediction result.

[0107] In step 7, inputting the final role embedding representation into a multi-layer perceptron for fusion update and using the span predictor to make a prediction to obtain a prediction result specifically includes the following steps:

[0108] Extract the role embedding from the prompt embedding containing the discourse context semantics and event pattern semantics to obtain the embedding representation of the role extracted from the prompt embedding;

[0109] Input the embedding representation of the role extracted from the prompt embedding and the query vector of the role in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion update to obtain the updated embedding of the role in the span start position predictor. The corresponding relationship in the process is as follows:

[0110] ;

[0111] Among them, represents the embedding of the role in the span start position predictor, represents the multi-layer perceptron, represents the prompt embedding containing the discourse context semantics and event pattern semantics in the th type of role embedding representation, Denotes the query vector of the role in the span start position predictor after rounds of iteration;

[0112] Input the embedding representation of the role extracted from the prompt embedding and the query vector of the role in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the role in the span end position predictor;

[0113] Input the passage embedding incorporating context semantics and the query vector of the word in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the passage in the span start position predictor. The corresponding relationship in the process is as follows:

[0114] ;

[0115] Among them, denotes the embedding of the passage in the span start position predictor ; Denotes the query vector of the passage in the span start position predictor after rounds of iteration;

[0116] Input the passage embedding incorporating context semantics and the query vector of the word in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the passage in the span end position predictor;

[0117] Respectively construct matrices for the updated embedding of the role in the span start position predictor, the updated embedding of the role in the span end position predictor, the updated embedding of the passage in the span start position predictor, and the updated embedding of the passage in the span end position predictor to obtain the learnable weight matrix of the role in the span start position predictor, the learnable weight matrix of the role in the span end position predictor, the learnable weight matrix of the passage in the span start position predictor, and the learnable weight matrix of the passage in the span end position predictor;

[0118] Use the learnable weight matrix of the role in the span start position predictor to perform a multiplication calculation on the updated embedding of the role in the span start position predictor. The corresponding relationship in the process is as follows:

[0119] ;

[0120] Among them, denotes the embedding of the role in the predicted span start position predictor,​​ represents the learnable weight matrix of the role in the span start position predictor during the prediction phase;

[0121] Using the constructed learnable weight matrix of the role in the span end position predictor, perform a multiplication calculation on the embedding of the role in the updated span end position predictor to obtain the role embedding in the predicted span end position predictor;

[0122] Using the constructed learnable weight matrix of the passage in the span start position predictor, perform a multiplication calculation on the embedding of the passage in the updated span start position predictor to obtain the passage embedding in the predicted span start position predictor;

[0123] Using the constructed learnable weight matrix of the passage in the span end position predictor, perform a multiplication calculation on the embedding of the passage in the updated span end position predictor to obtain the passage embedding in the predicted span end position predictor;

[0124] Using the role embedding in the predicted span start position predictor, perform a position prediction on the passage embedding in the predicted span start position predictor to obtain the distribution of the predicted span start for the role on the passage. The corresponding relationship in the process is as follows:

[0125] ;

[0126] where represents for the role the distribution of the predicted span start on the passage;

[0127] Using the role embedding in the predicted span end position predictor, perform a position prediction on the passage embedding in the predicted span end position predictor to obtain the distribution of the predicted span end for the role on the passage. The corresponding relationship in the process is as follows:

[0128] ;

[0129] where represents for the role the distribution of the predicted span end on the passage.

[0130] Based on the prediction results, construct a cross-entropy loss. The corresponding relationship in the process is as follows:

[0131] ;

[0132] where represents the probability distribution in the role span start position predictor, represents the role Probability distribution in the span end position predictor, representing the cross-entropy loss, representing the logarithmic function, representing the th character of a given event type, representing all the number of passages, representing the number of characters of a given event type, representing the number of documents.

[0133] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0134] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0135] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. A method for learning character and word representations driven by dual self-attention, characterized in that The method includes the following steps: Step 1: Construct a feature fusion module based on a linear transformation mechanism and a gated neural network, and construct a role-word self-attention module and a word-role self-attention module based on a bidirectional self-attention mechanism. The feature fusion module, the role-word self-attention module, the word-role self-attention module, a multi-layer perceptron, and a span predictor constitute a prediction model. Step 2: Input the text and provide the passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings. Step 3: Use the feature fusion module to capture the span correlation between the pattern semantics of the role and the context semantics of the word in different semantic embeddings, and obtain the initial embedding representation of the role and the initial embedding representation of the word. Step 4: Based on the role-word self-attention module, use the dual self-attention mechanism to calculate the attention weight of the role-word, and update the initial embedding representation of the role to obtain an updated role embedding representation. Based on the word-role self-attention module, use the dual self-attention mechanism to calculate the attention weight of the word-role, and update the initial embedding representation of the word to obtain an updated word embedding representation. Step 5: Perform interactive iteration on the updated role embedding representation and the updated word embedding representation to obtain the role embedding representation after the current interactive iteration and the word embedding representation after the current interactive iteration, respectively. Step 6: Repeat Step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation. Step 7: Input the final role embedding representation into the multi-layer perceptron for fusion update, and use the span predictor for prediction to obtain a prediction result. Construct a cross-entropy loss based on the prediction result, use the cross loss to optimize the prediction model to obtain an optimized prediction model, and use the optimized prediction model to obtain the final prediction result.

2. The method for learning character and word representations driven by dual self-attention according to claim 1, wherein In Step 2, input the text and provide the passage, perform syntactic and semantic analysis on the passage, and extract multi-level features in combination with a pre-trained language model to obtain different semantic embeddings, which specifically includes the following steps: Input the text and provide the passage, perform word segmentation analysis and syntactic analysis on the passage using a syntactic analysis tool in sequence, and perform word segmentation processing and random initialization operation on the passage using a pre-trained language model in sequence to obtain the embedding of the trigger word, the embedding of the word semantics, the part-of-speech embedding, the dependency relation type embedding, and the event type embedding. The relational expressions existing in the corresponding processes are as follows: ; Among them, represents the embedding representation of the trigger word, represents the pre-trained language model, represents the trigger word in the passage, represents the th word 's semantic embedding, represents the part-of-speech embedding, represents the random initialization operation, represents the dependency relation type embedding, represents the event type embedding, represents the role's embedding; Perform a marker insertion operation on the trigger word in the passage to obtain the passage after marker insertion, input the passage after marker insertion into the pre-trained language model for encoding through an encoder to obtain the output passage embedding, and decode the output passage embedding through a decoder to obtain the passage embedding incorporating context semantics. The relational expressions existing in the corresponding processes are as follows: ; Among them, represents the passage embedded representation output by the pre-trained language model encoder ; represents the passage after passing through the pre-trained language model encoder and the decoder and the generated embedded representation; Provide an event type hint for the passage, input the output passage embedding and the event type hint into the decoder for decoding to obtain a hint embedding containing the passage context semantics and the event pattern semantics. The relational expressions existing in the corresponding processes are as follows: ; Among them, represents a hint embedding containing the semantic meaning of the discourse context and the semantic meaning of the event pattern, represents the event type hint text.

3. The method for learning character and word representation driven by dual self-attention according to claim 2, wherein In step 3, the feature fusion module is used to capture the span correlation between the pattern semantics of different semantic embeddings for capturing roles and the context semantics of words, and obtain the initial embedding representations of roles and words. The specific steps are as follows: Input the embeddings of roles, event type embeddings, and trigger word embeddings into the feature fusion module, and perform vector fusion processing in combination with the span predictor to respectively obtain the query vector of the role in the span start position predictor and the query vector of the role in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the query vector of the role in the span start position predictor, represents the query vector of the role in the span end position predictor, represents being processed by the feature fusion module, represents the learnable weight matrix of the pattern semantics and context semantics in the span start position predictor, represents the learnable weight matrix of the pattern semantics and context semantics in the span end position predictor; Input the embeddings of word semantics, part-of-speech embeddings, and dependency relation type embeddings into the feature fusion module, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span start position predictor. Perform a linear transformation on the key vector of the word in the span start position predictor to obtain the value vector of the word in the span start position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span start position predictor, represents the value vector of the word in the span start position predictor, represents being processed by linear transformation; Input the embeddings of word semantics, part-of-speech embeddings, and dependency relation type embeddings into the feature fusion module again, and perform vector fusion processing in combination with the span predictor to obtain the key vector of the word in the span end position predictor. Perform a linear transformation on the key vector of the word in the span end position to obtain the value vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span end position predictor, represents the value vector of the word in the span end position predictor; Input the embeddings of word semantics, part-of-speech embeddings, and dependency relation type embeddings into the feature fusion module, and perform vector fusion processing in combination with the span predictor to respectively obtain the query vector of the word in the span start position predictor and the query vector of the word in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the query vector of the word in the span start position predictor, represents the query vector of the word in the span end position predictor; Input the embeddings of roles, event type embeddings, and trigger word embeddings into the feature fusion module, and perform vector fusion processing in combination with the span predictor to respectively obtain the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor. Perform linear transformations on the key vector of the role in the span start position predictor and the key vector of the role in the span end position predictor respectively to obtain the value vector of the role in the span start position predictor and the value vector of the role in the span end position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the role in the span start position predictor, represents the value vector of the role in the span start position predictor, represents the key vector of the role in the span end position predictor, represents the value vector of the role in the span end position predictor.

4. The method for learning character and word representations driven by dual self-attention according to claim 3, wherein In step 4, the binary self-attention mechanism is used to calculate the attention weights of role-words and update the initial embedding representation of the role to obtain the updated role embedding representation. The specific steps are as follows: Perform a multi-head self-attention operation on the query vector of the role in the span start position predictor in combination with the key vector of the word in the span start position predictor and the value vector of the word in the span start position predictor to obtain the attention weight matrix of the role in the span start position predictor. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the attention weight matrix of the role in the span start position predictor, represents the multi-head operation, represents being normalized, represents the passage all words in the matrix of the combination of key vectors in the span start position predictor, represents the transpose operation of the matrix, represents the passage all words in the matrix of the combination of value vectors in the span start position predictor, represents the key vector of the word in the span start position predictor the dimension of; Perform multi-head self-attention operation on the query vector of the role in the span end position predictor, combined with the key vector of the word in the span end position predictor and the value vector of the word in the span end position predictor, to obtain the attention weight matrix of the role in the span end position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the attention weight matrix of the role in the span end position predictor, represents the passage all words in the matrix of the combined key vectors of all words in the span end position predictor, represents the passage all words in the matrix of the combined value vectors of all words in the span end position predictor; Perform multiplication operation and non-linear activation processing on the attention weight matrix of the role in the span start position predictor and the value vector of the word in the span start position predictor in sequence, to obtain the updated query vector of the role in the span start position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the query vector of the updated role in the span start position predictor, indicates being processed by non-linear activation; Perform multiplication operation and non-linear activation processing on the attention weight matrix of the role in the span end position predictor and the value vector of the word in the span end position predictor in sequence, to obtain the updated query vector of the role in the span end position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the query vector of the updated role in the span end position predictor.

5. The method for learning character and word representation based on dual self-attention driving according to claim 4, characterized in that Use the dual self-attention mechanism to calculate the word-role attention weights and update the initial embedding representation of the word to obtain the updated word embedding representation. The specific steps are as follows: Perform multi-head self-attention operation on the query vector of the word in the span start position predictor, combined with the key vector of the role in the span start position predictor and the value vector of the role in the span start position predictor, to obtain the attention weight matrix of the word in the span start position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the attention weight matrix of the word in the span start position predictor, represents the role in the span start position predictor of the key vector represents the specified event type for all roles in the span start position predictor represents the specified event type for all roles in the span start position predictor Perform multi-head self-attention operation on the query vector of the word in the span end position predictor, combined with the key vector of the role in the span end position predictor and the value vector of the role in the span end position predictor, to obtain the attention weight matrix of the word in the span end position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the attention weight matrix of the word in the span end position predictor, represents the specified event type all the roles on the matrix of the combination of key vectors in the span end position predictor, represents the specified event type all the roles on the matrix of the combination of value vectors in the span end position predictor; Perform multiplication calculation and non-linear activation processing on the attention weight matrix of the word in the span start position predictor and the value vector of the role in the span start position predictor in sequence, to obtain the updated query vector of the word in the span start position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the query vector of the updated word in the span start position predictor; Perform multiplication calculation and non-linear activation processing on the attention weight matrix of the word in the span end position predictor and the value vector of the role in the span end position predictor in sequence, to obtain the updated query vector of the word in the span end position predictor. The relational expressions for the corresponding process are as follows: ; Among them, represents the query vector of the updated word in the span end position predictor.

6. The method for representation learning of roles and words based on dual self-attention driving according to claim 5, characterized in that In step 5, perform interactive iteration on the updated role embedding representation and the updated word embedding representation to obtain the current iteratively interactive role embedding representation and the current iteratively interactive word embedding representation respectively. The specific steps are as follows: Perform linear transformation processing on the updated query vector of the role in the span start position predictor and the updated query vector of the word in the span start position predictor respectively, to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction; Perform linear transformation processing on the query vector of the updated role in the span end position predictor and the query vector of the updated word in the span end position predictor respectively to obtain the query vector of the role in the span end position predictor after iterative interaction and the query vector of the word in the span end position predictor after iterative interaction; Use the query vector of the role in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span start position predictor to obtain the key vector of the role in the span start position predictor after interactive iteration; Use the query vector of the role in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span end position predictor to obtain the key vector of the role in the span end position predictor after interactive iteration; Use the query vector of the word in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span start position predictor to obtain the key vector of the word in the span start position predictor after interactive iteration; Use the query vector of the word in the span end position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span end position predictor to obtain the key vector of the word in the span end position predictor after interactive iteration.

7. The method for learning character and word representations driven by dual self-attention according to claim 6, wherein Perform linear transformation processing on the query vector of the updated role in the span start position predictor and the query vector of the updated word in the span start position predictor respectively to obtain the query vector of the role in the span start position predictor after iterative interaction and the query vector of the word in the span start position predictor after iterative interaction. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the query vector of the role in the span start position predictor after the -th iteration, represents the query vector of the word in the span start position predictor after the -th iteration, represents the query vector of the role in the span start position predictor after the -th iteration, represents the query vector of the word in the span start position predictor after the -th iteration; In the step of using the query vector of the role in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the role in the span start position predictor to obtain the key vector of the role in the span start position predictor after interactive iteration, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span start position predictor after the -th round of iteration; In the step of using the query vector of the word in the span start position predictor after iterative interaction to perform interactive iteration on the key vector of the word in the span start position predictor to obtain the key vector of the word in the span start position predictor after interactive iteration, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the key vector of the word in the span end position predictor after the round of iteration.

8. The method for learning the representation of roles and words based on dual self-attention driving according to claim 7, wherein In step 6, repeat step 5 in an iterative manner to obtain the final role embedding representation and the final word embedding representation, which specifically includes the following steps: Use the key vector of the role in the span start position predictor after interactive iteration as the input in the next round, and perform interactive iteration on the query vector of the role in the span start position predictor after interactive iteration in an iterative manner to obtain the query vector of the role in the span start position predictor after the next round of interactive iteration; Use the key vector of the role in the span end position predictor after interactive iteration as the input in the next round, and perform interactive iteration on the query vector of the role in the span end position predictor after interactive iteration in an iterative manner to obtain the query vector of the role in the span end position predictor after the next round of interactive iteration; Using the key vector of the word after interactive iteration in the span start position predictor as input in the next round, interactively iterate the query vector of the word after interactive iteration in the span start position predictor in an iterative manner to obtain the query vector of the word after the next round of interactive iteration in the span start position predictor; Using the key vector of the word after interactive iteration in the span end position predictor as input in the next round, interactively iterate the query vector of the word after interactive iteration in the span end position predictor in an iterative manner to obtain the query vector of the word after the next round of interactive iteration in the span end position predictor.

9. The method for learning character and word representations driven by dual self-attention according to claim 8, wherein In step 7, input the final role embedding representation into a multi-layer perceptron for fusion and update, and use the span predictor to make a prediction to obtain the prediction result, which specifically includes the following steps: Extract the role embedding from the hint embedding containing the context semantics of the passage and the event pattern semantics to obtain the embedding representation of the role extracted from the hint embedding; Input the embedding representation of the role extracted from the hint embedding and the query vector of the role in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the role in the span start position predictor. The corresponding relationship in the process is as follows: ; Among them, represents the embedding of the role in the span start position predictor ; represents a multi-layer perceptron ; represents the hint embedding containing the semantic information of the discourse context and the event pattern in the -th role embedding representation ; represents the query vector of the role in the span start position predictor after rounds of iteration Input the embedding representation of the role extracted from the hint embedding and the query vector of the role in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the role in the span end position predictor; Input the passage embedding incorporating the context semantics and the query vector of the word in the span start position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the passage in the span start position predictor. The corresponding relationship in the process is as follows: ; Among them, represents the embedding of the passage in the span start position predictor ; represents the query vector of the passage in the span start position predictor after rounds of iteration; ​ Input the passage embedding incorporating the context semantics and the query vector of the word in the span end position predictor after the next round of interactive iteration into a multi-layer perceptron for fusion and update to obtain the updated embedding of the passage in the span end position predictor; Respectively construct matrices for the updated embedding of the role in the span start position predictor, the updated embedding of the role in the span end position predictor, the updated embedding of the passage in the span start position predictor, and the updated embedding of the passage in the span end position predictor to obtain the learnable weight matrix of the role in the span start position predictor, the learnable weight matrix of the role in the span end position predictor, the learnable weight matrix of the passage in the span start position predictor, and the learnable weight matrix of the passage in the span end position predictor; Use the learnable weight matrix of the role in the span start position predictor to perform a multiplication calculation on the updated embedding of the role in the span start position predictor to obtain the role embedding in the predicted span start position predictor. The corresponding relationship in the process is as follows: ; Among them, represents the role embedding in the predictor for the start position of the prediction span, represents the learnable weight matrix of the role in the prediction phase for the start position of the span in the predictor; Use the learnable weight matrix of the role in the span end position predictor to perform a multiplication calculation on the updated embedding of the role in the span end position predictor to obtain the role embedding in the predicted span end position predictor; Using the learnable weight matrix of the passage in the span start position predictor, perform a multiplication calculation on the passage embedding in the updated span start position predictor to obtain the passage embedding in the predicted span start position predictor; Using the learnable weight matrix of the passage in the span end position predictor, perform a multiplication calculation on the passage embedding in the updated span end position predictor to obtain the passage embedding in the predicted span end position predictor; Using the role embedding in the predicted span start position predictor, perform a position prediction on the passage embedding in the predicted span start position predictor to obtain the distribution of the predicted span start for the role on the passage. The corresponding relationship in the process is as follows: ; Among them, represents the distribution starting from the span predicted for the role in the passage ; Using the role embedding in the predicted span end position predictor, perform a position prediction on the passage embedding in the predicted span end position predictor to obtain the distribution of the predicted span end for the role on the passage. The corresponding relationship in the process is as follows: ; Among them, represents the distribution of the end of the span predicted for the role in the passage above.

10. The method for learning character and word representations driven by dual self-attention according to claim 9, wherein Construct a cross-entropy loss based on the prediction results. The corresponding relationship in the process is as follows: ; Among them, represents the probability distribution in the span start position predictor for the role represents the probability distribution in the span end position predictor for the role represents the cross-entropy loss represents the logarithmic function represents the th role for a given event type represents all the number of passages represents the number of roles for a given event type represents the number of documents​​

Citation Information

Patent Citations

  • Method for identifying chapter-level event roles based on knowledge graph information guidance

    CN114880434A

  • Relying on discourse analysis to answer complex questions by neural machine reading comprehension

    US20220138432A1