Character semantics capture method and system based on dynamic event prompts

By introducing dynamic event prompts and role interaction modules into the event argument extraction method, the problems of insufficient universality of event prompt templates and low template construction efficiency in the existing technology are solved, and the more efficient and semantic rich event argument extraction effect is achieved.

CN119783683BActive Publication Date: 2025-05-23JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510284611.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-23
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing prompt-based event argument extraction method has limitations, including sharing event prompt templates with different event types, ignoring the differences in content between documents, high cost and low efficiency in manual construction of templates, and low performance in automatic model construction templates.

Method used

A role semantic capture method based on dynamic event prompts is proposed. By preprocessing and encoding the document, the event prompt embedding representation is obtained using the decoder, the role feature representation is obtained and average pooled, and the role embedding representation is obtained. Then, the training parameters are introduced, the role start span selector and the end span selector are obtained, and the candidate span score is calculated through the greedy search algorithm to obtain the reconstructed event prompt template. Based on the reconstructed template, the reconstructed role representation is obtained using the encoder, and the role interaction layer processing is used to determine the predicted span of the role and calculate the best matching result.

Benefits of technology

It realizes dynamic reconstruction of event prompt templates, adapting to different documents and event types, improving the efficiency and quality of template construction, reducing labor costs, enhancing the richness of role semantics, and allowing the model to more comprehensively understand the semantic connotation of events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783683B_ABST
    Figure CN119783683B_ABST
Patent Text Reader

Abstract

The present invention proposes a role semantics capture method and system based on dynamic event prompts, the method comprising: encoding a preprocessed document to obtain an encoded document representation; obtaining an event prompt embedding representation using an event prompt board and the encoded document representation; obtaining a role embedding representation based on the event prompt embedding representation; obtaining a role start span selector and an end span selector respectively according to the role embedding representation; obtaining a candidate span of the role using the encoded document representation and the role start span selector and the end span selector; obtaining a reconstructed event prompt template through the candidate span of the role; obtaining a role representation processed by a role interaction layer based on the reconstructed event prompt template; training a model based on the role representation processed by the role interaction layer to obtain an argument extraction result. The present invention enhances the richness of role semantics through a role interaction module, so that the model can understand the semantic connotation of the event more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information extraction, and in particular to a method and system for capturing character semantics based on dynamic event prompts. Background Art

[0002] In natural language processing, event argument extraction is a key task that aims to identify events and their related arguments from text. Existing prompt-based event argument extraction methods mainly rely on preset prompt information and identify event arguments in input text by modeling role information in event prompt templates. These methods have the following limitations: different event types share event prompt templates, ignoring the semantic uniqueness of various events; different documents have the same prompt template content, failing to fully consider the differences in content between documents; manual template construction is costly and inefficient, and the performance of automatic template construction by the model is low. Summary of the invention

[0003] In view of the above situation, the main purpose of the present invention is to propose a character semantics capture method based on dynamic event prompts to solve the above technical problems.

[0004] The present invention proposes a character semantics capture method based on dynamic event prompts, the method comprising the following steps:

[0005] Step 1: Preprocess the document and then encode it to obtain the encoded document representation; use the original event prompt board and the encoded document representation to obtain the event prompt embedding representation through the decoder; obtain the character feature representation based on the event prompt embedding representation; obtain the character embedding representation by average pooling the character feature representation;

[0006] Step 2: introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation;

[0007] Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively;

[0008] The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained;

[0009] The reconstructed event prompt template is obtained through the candidate span of the role;

[0010] Step 3: Based on the reconstructed event prompt template, the encoder is used to obtain the reconstructed role representation, and then the predicted event instance representation is determined by the reconstructed role representation;

[0011] Using the predicted event instance representation and the gating mechanism, we can calculate the embedded representation of the predicted event instance after interaction.

[0012] Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer;

[0013] Step 4: Determine the predicted span of the role through the role representation containing local document information, then calculate the best matching result using the actual span of the role and the predicted span of the role, and then further obtain the probability distribution of the role's starting position;

[0014] The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model;

[0015] The trained model is used to obtain the argument extraction results.

[0016] The present invention also proposes a character semantics capture system based on dynamic event prompts, the system comprising:

[0017] Semantic encoding module, used to:

[0018] The document is preprocessed and then encoded to obtain an encoded document representation; the original event prompt board and the encoded document representation are used to obtain an event prompt embedding representation through a decoder; the character feature representation is obtained based on the event prompt embedding representation; the character feature representation is averaged and pooled to obtain the character embedding representation;

[0019] Time prompts dynamically build modules for:

[0020] Introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation;

[0021] Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively;

[0022] The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained;

[0023] The reconstructed event prompt template is obtained through the candidate span of the role;

[0024] Role interaction module, used to:

[0025] Based on the reconstructed event prompt template, use the encoder to obtain the reconstructed role representation, and then determine the predicted event instance representation through the reconstructed role representation;

[0026] Use the predicted event instance representation, and combine with the gating mechanism to calculate the embedded representation after the interaction of the predicted event instance;

[0027] Based on the embedded representation after the interaction of the predicted event instance and the reconstructed role representation, through the role interaction layer, obtain the role representation processed by the role interaction layer;

[0028] The event argument extraction module is used for:

[0029] Determine the predicted span of the role through the role representation containing local document information, then calculate the best matching result using the actual span of the role and the predicted span of the role, and then further obtain the probability distribution of the start position of the role;

[0030] Construct a cross-entropy loss function through the actual span of the role, and then use the cross-entropy loss function and the probability distribution of the start position of the role to train the model to obtain the trained model;

[0031] Use the trained model to obtain the argument extraction result.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] 1. The present invention proposes a strategy for dynamically reconstructing event prompts based on predicted event instances. This strategy is a general framework for constructing document prompt templates, which can match the document context, adapt to different documents and event types. At the same time, the semi-automatic template construction method combines human experience and machine learning technology, improving the efficiency and quality of template construction and reducing labor costs;

[0034] 2. The present invention enhances the richness of role semantics by designing a role interaction module, enabling the model to more comprehensively understand the semantic connotation of events;

[0035] 3. The present invention develops a document-level event argument extraction model solution based on the strategy of dynamically reconstructing event prompts based on predicted event instances and the role interaction module.

[0036] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. Brief Description of the Drawings

[0037] Figure 1 It is a flowchart of the method for capturing role semantics based on dynamic event prompts proposed by the present invention;

[0038] Figure 2 This is a framework diagram of the role semantics capture system based on dynamic event prompts proposed by the present invention. DETAILED DESCRIPTION

[0039] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0040] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0041] See also Figure 1 and Figure 2 The embodiment of the present invention proposes a character semantics capture method based on dynamic event prompts, which includes the following steps:

[0042] Step 1: Preprocess the document and then encode it to obtain the encoded document representation; use the original event prompt board and the encoded document representation to obtain the event prompt embedding representation through the decoder; obtain the character feature representation based on the event prompt embedding representation; obtain the character embedding representation by average pooling the character feature representation;

[0043] In step 1, the document is preprocessed and then encoded to obtain the encoded document representation; the event prompt embedding representation is obtained by the decoder using the original event prompt board and the encoded document representation; the character feature representation is obtained based on the event prompt embedding representation; the character feature representation is averaged and pooled to obtain the character embedding representation. The specific steps are as follows:

[0044] Perform special marking processing on the document to obtain the pre-processed document;

[0045] The preprocessed document is encoded to obtain the encoded document representation. The relationship between the corresponding process is:

[0046] ;

[0047] in, Represents the encoded document representation, It means that it has been processed by the encoder of the BART model. Represents the preprocessed document;

[0048] After splicing the event information in the original event prompt board, insert a special mark to obtain the event prompt board;

[0049] Using the event prompt board and the encoded document representation, the event prompt embedding representation is obtained through the decoder. The relationship between the corresponding process is:

[0050] ;

[0051] in, Indicates that the event prompt is embedded in the representation, It means that it has been processed by the decoder of the BART model. Indicates event prompt board;

[0052] According to the slots of each character in the event prompt board, the characteristic representation of the character is obtained by intercepting the position;

[0053] By performing average pooling on the feature representation of the role, the role embedding representation is obtained. The relationship between the corresponding process is:

[0054] ;

[0055] in, Indicates The embedding representation of the characters, It means that after the average pooling operation, Indicates Characteristic representation of a role.

[0056] Step 2: introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation;

[0057] Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively;

[0058] The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained;

[0059] The reconstructed event prompt template is obtained through the candidate span of the role;

[0060] In step 2, the training parameters are introduced, and the role start span selector and the end span selector are obtained according to the role embedding representation. The corresponding relationship is:

[0061] ;

[0062] in, Indicates the start span selector for the th character, Indicates the end span selector for the th character, and both represent training parameters, Indicates element-wise multiplication;

[0063] Using the encoded document representation, the start span selector, and the end span selector for the character, the start boundary distribution and the end boundary distribution of the character in the input text are obtained respectively. The relational expressions for the corresponding processes are:

[0064] ;

[0065] Among them, Indicates the start boundary distribution of the th character in the input text, Indicates the end boundary distribution of the th character in the input text;

[0066] Based on the start boundary distribution and the end boundary distribution of the character in the input text, the occurrence of the character is predicted, and then all candidate span scores of the character are calculated through the greedy search algorithm. Subsequently, the candidate spans of the character are further obtained. The specific steps are as follows:

[0067] Based on the start boundary distribution and the end boundary distribution of the character in the input text, the occurrence of the character is predicted, and then all candidate span scores of the character are calculated through the greedy search algorithm. The relational expressions for the corresponding processes are:

[0068] ;

[0069] Among them, Indicates all candidate span scores of the th character, Indicates the start position of the candidate span, Indicates the end position of the candidate span, Indicates the log probability of the th character at the start position, Indicates the log probability of the th character at the end position;

[0070] Select the combination with the highest score among all candidate span scores of the character as the candidate span of the character. The relational expressions for the corresponding processes are:

[0071] ;

[0072] Among them, Indicates the The candidate span of roles, represents the candidate span set, Indicates The starting span of each character, Indicates The ending span of a character.

[0073] Furthermore, according to the candidate span of the obtained role, the predicted event instance is obtained in the preprocessed document by intercepting by position. The predicted event instance obtained by predicting the starting position and the ending position of the event instance is filled into the corresponding position of the document prompt template to form a new event prompt template, and the predicted event instance of the next role is identified until all roles are identified to obtain the reconstructed event prompt template.

[0074] Step 3: Based on the reconstructed event prompt template, the encoder is used to obtain the reconstructed role representation, and then the predicted event instance representation is determined by the reconstructed role representation;

[0075] Using the predicted event instance representation and the gating mechanism, we can calculate the embedded representation of the predicted event instance after interaction.

[0076] Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer;

[0077] In step 3, the predicted event instance representation is used in combination with the gating mechanism to calculate the embedded representation of the predicted event instance after interaction. The specific steps are as follows:

[0078] Receive the predicted event instance representation and the hidden state of the previous moment, map them to a value between 0 and 1 through the Sigmoid activation function, and get the output of the forget gate. The corresponding process has the following relationship:

[0079] ;

[0080] in, represents the output of the forget gate, Indicates that it has been processed by the Sigmoid activation function. Indicates the hidden state at the previous moment, represents the predicted event instance representation, represents the bias vector, represents the weight matrix;

[0081] Receive the predicted event instance representation and the hidden state of the previous moment, obtain the update ratio through the Sigmoid activation function, and transform it through the Tanh activation function to obtain the candidate cell state. The relationship between the corresponding process is:

[0082] ;

[0083] in, represents the candidate cell state, Indicates that it has been processed by the Tanh activation function. and Both represent weight matrices, and Both represent bias vectors, Indicates the update ratio;

[0084] The candidate cell state is updated using the update ratio and the output of the forget gate to obtain the cell state at the current moment. The corresponding process has the following relationship:

[0085] ;

[0086] in, Indicates the cell state at the current moment, Indicates the cell state at the previous moment;

[0087] Based on the predicted event instance representation, the hidden state at the previous moment, and the cell state at the current moment, the embedded representation after the predicted event instance interaction is obtained. The relationship between the corresponding process is:

[0088] ;

[0089] in, represents the output of the output gate, represents the weight matrix, represents the bias vector, Represents the embedding representation after the interaction of the predicted event instance;

[0090] Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer. The specific steps are as follows:

[0091] By fusing the reconstructed role representation and the embedded representation after the interaction of the predicted event instance, we can obtain the role representation containing local document information. The relationship between the corresponding process is:

[0092] ;

[0093] in, Indicates the first The representation of a role, Indicates that after the fusion operation, Represents the reconstruction The representation of a role, Indicates The embedding representation of the predicted event instances of the roles after interaction;

[0094] The role representation containing local document information is passed through the role interaction layer and calculated by the LSTM network to obtain the role representation processed by the role interaction layer. The relationship between the corresponding process is:

[0095] ;

[0096] in, All represent the role representation processed by the role interaction layer. Indicates the calculation after LSTM network, Both represent role representations that contain local document information.

[0097] Step 4: Determine the predicted span of the role through the role representation containing local document information, then calculate the best matching result using the actual span of the role and the predicted span of the role, and then further obtain the probability distribution of the role's starting position;

[0098] The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model;

[0099] Use the trained model to obtain the argument extraction results;

[0100] In step 4, the predicted span of the role is determined by the role representation containing local document information, and then the best matching result is calculated using the actual span of the role and the predicted span of the role, and then the probability distribution of the starting position of the role is further obtained. The specific steps are as follows:

[0101] Using the role representation processed by the role interaction layer to perform predictive argument recognition, we can obtain the starting boundary distribution of the start span selector and the ending boundary distribution of the end span selector. The relationship between the corresponding processes is:

[0102] ;

[0103] in, Represents the start span selector, Indicates the end span selector, and All represent trainable parameters, Indicates the starting boundary distribution of the start span selector. Indicates the ending boundary distribution of the ending span selector;

[0104] Based on the starting boundary distribution of the start span selector and the ending boundary distribution of the end span selector, the span position of the candidate argument is predicted to obtain the predicted span of the role. The corresponding process has the following relationship:

[0105] ;

[0106] in, Indicates The prediction span of each character, Indicates the starting position of the candidate span, Indicates the end position of the candidate span, Indicates The logarithmic probability of a character in the starting position, Indicates The logarithmic probability of a character in the final position, represents the candidate span set;

[0107] Using the actual span of the role and the predicted span of the role, the best matching result is obtained. The relationship between the corresponding process is:

[0108] ;

[0109] in, represents the best matching result, Indicates the minimum value between the number of appearances of the character in the prompt template and the number of all characters in the time prompt. It means that after L1 norm processing, Indicates The actual starting span of each character, Indicates The actual ending span of each character, represents the permutation distribution, It indicates the index of the element. Indicates the order of The starting span of the element corresponding to each role, Indicates the order of The ending span of the element corresponding to each role;

[0110] The probability distribution of the character's start and end positions is calculated based on the best matching result. The corresponding relationship is:

[0111] ;

[0112] in, Indicates The probability distribution of the starting positions of the characters, Indicates The probability distribution of the ending position of each character, Indicates the optimal match after The starting boundary distribution of each character, Indicates the optimal match after The ending boundary distribution of each character;

[0113] The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model. The relationship of the cross entropy loss function is:

[0114] ;

[0115] in, represents the cross entropy loss function, Indicates event prompts. represents the probability distribution of the starting position predicted by the model, represents the probability distribution of the end position predicted by the model, Indicates The starting position of the actual argument of each role, Indicates The end position of the actual argument of each role, Indicates that the value range is all contexts in the data set. Indicates the number of all roles for event prompts.

[0116] Furthermore, the argument recognition F1 value and argument classification F1 value are used as evaluation indicators to evaluate the argument extraction results. The calculation formula is:

[0117] ;

[0118] in, Indicates the accuracy, represents the recall rate, represents the number of samples predicted to be positive and whose true value is positive, Represents the number of samples that are predicted to be positive but the true value is negative. Represents the number of samples that are predicted to be negative but the true value is positive. Represents the harmonic mean.

[0119] See also Figure 2 The embodiment of the present invention further provides a role semantics capture system based on dynamic event prompts, the system comprising:

[0120] Semantic encoding module, used to:

[0121] The document is preprocessed and then encoded to obtain an encoded document representation; the original event prompt board and the encoded document representation are used to obtain an event prompt embedding representation through a decoder; the character feature representation is obtained based on the event prompt embedding representation; the character feature representation is averaged and pooled to obtain the character embedding representation;

[0122] Time prompts dynamically build modules for:

[0123] Introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation;

[0124] Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively;

[0125] The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained;

[0126] The reconstructed event prompt template is obtained through the candidate span of the role;

[0127] Role interaction module, used to:

[0128] Based on the reconstructed event prompt template, the encoder is used to obtain the reconstructed role representation, and then the predicted event instance representation is determined by the reconstructed role representation;

[0129] Using the predicted event instance representation and the gating mechanism, we can calculate the embedded representation of the predicted event instance after interaction.

[0130] Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer;

[0131] Event argument extraction module, used for:

[0132] The predicted span of the role is determined by the role representation containing local document information, and then the best matching result is calculated using the actual span of the role and the predicted span of the role, and then the probability distribution of the starting position of the role is further obtained;

[0133] The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model;

[0134] The trained model is used to obtain the argument extraction results.

[0135] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0136] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0137] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A character semantics capture method based on dynamic event prompts, characterized in that: The method comprises the following steps: Step 1: Preprocess the document and then encode it to obtain the encoded document representation; use the original event prompt board and the encoded document representation to obtain the event prompt embedding representation through the decoder; obtain the character feature representation based on the event prompt embedding representation; obtain the character embedding representation by average pooling the character feature representation; Step 2: Introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation; Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively; The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained; The reconstructed event prompt template is obtained through the candidate span of the role; The steps of obtaining the reconstructed event prompt template through the candidate span of the role are as follows: according to the candidate span of the obtained role, the predicted event instance is obtained in the preprocessed document by intercepting by position; the predicted event instance obtained by the starting position and the ending position of the predicted event instance is filled into the corresponding position of the document prompt template to form a new event prompt template, and the predicted event instance of the next role is identified until all roles are identified to obtain the reconstructed event prompt template; Step 3: Based on the reconstructed event prompt template, the encoder is used to obtain the reconstructed role representation, and then the predicted event instance representation is determined by the reconstructed role representation; Using the predicted event instance representation and the gating mechanism, we can calculate the embedded representation of the predicted event instance after interaction. Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer; Step 4: Determine the predicted span of the role through the role representation containing local document information, then calculate the best matching result using the actual span of the role and the predicted span of the role, and then further obtain the probability distribution of the role's starting position; The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model; The trained model is used to obtain the argument extraction results.

2. The character semantics capture method based on dynamic event prompts according to claim 1 is characterized in that: In the step 1, the document is preprocessed and then encoded to obtain an encoded document representation; the event prompt embedding representation is obtained by a decoder using the original event prompt board and the encoded document representation; the character feature representation is obtained according to the event prompt embedding representation; the character feature representation is average pooled to obtain the character embedding representation, and the specific steps are as follows: Perform special marking processing on the document to obtain the pre-processed document; The preprocessed document is encoded to obtain the encoded document representation. The relationship between the corresponding process is: ; in, Represents the encoded document representation, It means that it has been processed by the encoder of the BART model. Represents the preprocessed document; After splicing the event information in the original event prompt board, insert a special mark to obtain the event prompt board; Using the event prompt board and the encoded document representation, the event prompt embedding representation is obtained through the decoder. The corresponding process has the following relationship: ; in, Indicates that the event prompt is embedded in the representation, It means that it has been processed by the decoder of the BART model. Indicates event prompt board; According to the slots of each character in the event prompt board, the characteristic representation of the character is obtained by intercepting the position; By performing average pooling on the feature representation of the role, the role embedding representation is obtained. The relationship between the corresponding process is: ; in, Indicates The embedding representation of the characters, It means that after the average pooling operation, Indicates Characteristic representation of a role.

3. The character semantics capture method based on dynamic event prompts according to claim 2 is characterized in that: In step 2, training parameters are introduced, and the role start span selector and the end span selector are obtained according to the role embedding representation. The relationship between the corresponding process is: ; in, Indicates The starting span selector for a role, Indicates The ending span selector for the character, and Both represent training parameters, Represents element-wise multiplication.

4. The character semantics capture method based on dynamic event prompts according to claim 3 is characterized in that: In step 2, the encoded document representation, the role start span selector and the end span selector are used to obtain the start boundary distribution and the end boundary distribution of the role in the input text respectively. The relationship between the corresponding processes is: ; in, Indicates The distribution of the starting boundaries of characters in the input text, Indicates The ending boundary distribution of characters in the input text.

5. The character semantics capture method based on dynamic event prompts according to claim 4 is characterized in that: In step 2, the appearance of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then all candidate span scores of the role are calculated by a greedy search algorithm, and then the candidate span of the role is further obtained. The specific steps are as follows: The appearance of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm. The relationship between the corresponding process is: ; in, Indicates All candidate span scores for a role, Indicates the starting position of the candidate span, Indicates the end position of the candidate span, Indicates The logarithmic probability of a character in the starting position, Indicates The logarithmic probability of a character being in the final position; The combination with the highest score among all candidate span scores of the role is selected as the candidate span of the role. The relationship between the corresponding process is: ; in, Indicates The candidate span of roles, represents the candidate span set, Indicates The starting span of each character, Indicates The ending span of a character.

6. The character semantics capture method based on dynamic event prompts according to claim 5 is characterized in that: In step 3, the predicted event instance representation is used in combination with the gating mechanism to calculate the embedded representation of the predicted event instance after interaction. The specific steps are as follows: Receive the predicted event instance representation and the hidden state of the previous moment, map them to a value between 0 and 1 through the Sigmoid activation function, and get the output of the forget gate. The corresponding process has the following relationship: ; in, represents the output of the forget gate, Indicates that it has been processed by the Sigmoid activation function. Indicates the hidden state at the previous moment, represents the predicted event instance representation, represents the bias vector, represents the weight matrix; Receive the predicted event instance representation and the hidden state of the previous moment, obtain the update ratio through the Sigmoid activation function, and transform it through the Tanh activation function to obtain the candidate cell state. The relationship between the corresponding process is: ; in, represents the candidate cell state, Indicates that it has been processed by the Tanh activation function. and Both represent weight matrices, and Both represent bias vectors, Indicates the update ratio; The candidate cell state is updated using the update ratio and the output of the forget gate to obtain the cell state at the current moment. The corresponding process has the following relationship: ; in, Indicates the cell state at the current moment, Indicates the cell state at the previous moment; Based on the predicted event instance representation, the hidden state at the previous moment, and the cell state at the current moment, the embedded representation after the predicted event instance interaction is obtained. The relationship between the corresponding process is: ; in, represents the output of the output gate, represents the weight matrix, represents the bias vector, Represents the embedding representation after the interaction of the predicted event instance.

7. The character semantics capture method based on dynamic event prompts according to claim 6 is characterized in that: In step 3, based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer. The specific steps are as follows: By fusing the reconstructed role representation and the embedded representation after the interaction of the predicted event instance, we can obtain the role representation containing local document information. The relationship between the corresponding process is: ; in, Indicates the first The representation of a role, Indicates that after the fusion operation, Represents the reconstruction The representation of a role, Indicates The embedding representation of the predicted event instances of the roles after interaction; The role representation containing local document information is passed through the role interaction layer and calculated by the LSTM network to obtain the role representation processed by the role interaction layer. The relationship between the corresponding process is: ; in, All represent the role representation processed by the role interaction layer. Indicates the calculation after the LSTM network. Both represent role representations that contain local document information.

8. The character semantics capture method based on dynamic event prompts according to claim 7 is characterized in that: In step 4, the predicted span of the role is determined by the role representation containing the local document information, and then the best matching result is calculated by using the actual span of the role and the predicted span of the role, and then the probability distribution of the starting position of the role is further obtained. The specific steps are as follows: Using the role representation processed by the role interaction layer to perform predictive argument recognition, we can obtain the starting boundary distribution of the start span selector and the ending boundary distribution of the end span selector. The relationship between the corresponding processes is: ; in, Represents the start span selector, Indicates the end span selector, and All represent trainable parameters, Indicates the starting boundary distribution of the start span selector. Indicates the ending boundary distribution of the ending span selector; Based on the starting boundary distribution of the start span selector and the ending boundary distribution of the end span selector, the span position of the candidate argument is predicted to obtain the predicted span of the role. The corresponding process has the following relationship: ; in, Indicates The prediction span of each character, Indicates the starting position of the candidate span, Indicates the end position of the candidate span, Indicates The logarithmic probability of a character in the starting position, Indicates The logarithmic probability of a character in the final position, represents the candidate span set; Using the actual span of the role and the predicted span of the role, the best matching result is obtained. The relationship between the corresponding process is: ; in, represents the best matching result, Indicates the minimum value between the number of appearances of the character in the prompt template and the number of all characters in the time prompt. It means that after L1 norm processing, Indicates The actual starting span of each character, Indicates The actual ending span of each character, represents the permutation distribution, It indicates the index of the element. Indicates the order of The starting span of the element corresponding to each role, Indicates the order of The ending span of the element corresponding to each role; The probability distribution of the character's start and end positions is calculated based on the best matching result. The corresponding relationship is: ; in, Indicates The probability distribution of the starting positions of the characters, Indicates The probability distribution of the ending position of each character, Indicates the optimal match after The starting boundary distribution of each character, Indicates the optimal match after The ending boundary distribution of each character.

9. The character semantics capture method based on dynamic event prompts according to claim 8 is characterized in that: In step 4, a cross entropy loss function is constructed by the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain a trained model, wherein the relationship of the cross entropy loss function is: ; in, represents the cross entropy loss function, Indicates event prompts. represents the probability distribution of the starting position predicted by the model, represents the probability distribution of the end position predicted by the model, Indicates The starting position of the actual argument of each character, Indicates The end position of the actual argument of the role, Indicates that the value range is all contexts in the data set. Indicates the number of all roles for event prompts.

10. A character semantics capture system based on dynamic event prompts, characterized in that: The system applies any one of claims 1 to 9 of the character semantics capture method based on dynamic event prompts, and the system comprises: Semantic encoding module, used to: The document is preprocessed and then encoded to obtain an encoded document representation; the original event prompt board and the encoded document representation are used to obtain an event prompt embedding representation through a decoder; the character feature representation is obtained based on the event prompt embedding representation; the character feature representation is averaged and pooled to obtain the character embedding representation; Time prompts dynamically build modules for: Introduce training parameters and obtain the role start span selector and end span selector according to the role embedding representation; Using the encoded document representation, the role start span selector and the end span selector, the start boundary distribution and the end boundary distribution of the role in the input text are obtained respectively; The occurrence of the role is predicted based on the distribution of the starting boundary and the ending boundary of the role in the input text, and then the scores of all candidate spans of the role are calculated by the greedy search algorithm, and then the candidate spans of the role are further obtained; The reconstructed event prompt template is obtained through the candidate span of the role; The steps of obtaining the reconstructed event prompt template through the candidate span of the role are as follows: according to the candidate span of the obtained role, the predicted event instance is obtained in the preprocessed document by intercepting by position; the predicted event instance obtained by the starting position and the ending position of the predicted event instance is filled into the corresponding position of the document prompt template to form a new event prompt template, and the predicted event instance of the next role is identified until all roles are identified to obtain the reconstructed event prompt template; Role interaction module, used to: Based on the reconstructed event prompt template, the encoder is used to obtain the reconstructed role representation, and then the predicted event instance representation is determined by the reconstructed role representation; Using the predicted event instance representation and the gating mechanism, we can calculate the embedded representation of the predicted event instance after interaction. Based on the embedded representation after the predicted event instance interaction and the reconstructed role representation, the role representation processed by the role interaction layer is obtained through the role interaction layer; Event argument extraction module, used for: The predicted span of the role is determined by the role representation containing local document information, and then the best matching result is calculated using the actual span of the role and the predicted span of the role, and then the probability distribution of the starting position of the role is further obtained; The cross entropy loss function is constructed through the actual span of the character, and then the model is trained using the cross entropy loss function and the probability distribution of the character's starting position to obtain the trained model; The trained model is used to obtain the argument extraction results.

Citation Information

Patent Citations

  • Text extraction method and device and electronic equipment

    CN117634449A

  • Document-level event argument extraction method and device, equipment and medium

    CN118673899A