Financial Document-Level Event Extraction Method and System Based on Multi-Semantic Enhancement

By adopting multiple semantic enhancement methods in financial document-level event extraction, the shortcomings of the prior art in processing noise data, complex financial terms and cross-sentence and cross-paragraph information are solved, and higher event extraction accuracy and efficiency are achieved.

CN119808793BActive Publication Date: 2025-06-17WUHAN LANSHAN TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286113.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-17
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Existing financial document-level event extraction methods perform poorly in processing noise data, complex financial terms, and cross-sentence and cross-paragraph information, resulting in inaccuracy and inefficiency of event extraction.

Method used

Using a method based on multiple semantic enhancement, financial documents are preprocessed, semantic matching optimization module, multi-grained semantic enhancement module and event center heterogeneous graph structure module, financial documents are preprocessed, semantic matching optimization, multi-grained semantic enhancement and heterogeneous graph structure analysis to extract event information.

Benefits of technology

Effectively filtering noise data improves the ability to understand complex financial terms, enhances the ability to capture information across sentences and paragraphs, and significantly improves the accuracy and efficiency of financial document-level event extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808793B_ABST
    Figure CN119808793B_ABST
Patent Text Reader

Abstract

The present invention provides a financial document-level event extraction method and system based on multi-semantic enhancement, belonging to the field of natural language processing, including: constructing a financial document event pattern, where the financial document event pattern includes an event type set and an argument role set of the event type; obtaining financial document data to be extracted, preprocessing and annotating the financial document data to be extracted based on the event type and argument role to obtain annotated financial document data; inputting the annotated financial document data into a trained financial document-level event extraction model, and using the semantic matching optimization module, multi-granularity semantic enhancement module and event center heterogeneous graph structure module of the financial document-level event extraction model to process the annotated financial document data to obtain an extraction result including event types and event arguments. The present invention realizes extracting event information from multi-level and cross-paragraph texts when processing financial documents, improving the accuracy and efficiency of event extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a financial document-level event extraction method and system based on multiple semantic enhancements. Background Art

[0002] Against the backdrop of the rapid development of information technology, event extraction technology in natural language processing has become an important research direction. Especially driven by financial technology, the financial industry has generated a large amount of unstructured text data, such as news reports, financial announcements, and analytical reports. Traditional event extraction methods are mostly based on rules and pattern matching. These methods rely on predefined event patterns. Although they can achieve good results in specific fields, they often perform poorly when dealing with new fields and complex texts. With the rise of deep learning, event extraction methods based on deep learning have gradually replaced traditional methods, especially the introduction of word embedding models (such as Word2Vec, BERT) and graph neural networks, which have greatly improved the model's ability to understand events and their semantic relationships.

[0003] Existing event extraction methods are mainly based on sentence-level extraction technology, which is mostly used to obtain event information from news headlines. However, financial events usually span multiple sentences or even multiple paragraphs, and a single sentence-level event extraction technology is difficult to capture the full picture of the event. For example, a company's financial report may contain multiple levels of information, including performance, financial indicators, business prospects, etc. Only by analyzing the entire report at the document level can relevant event information be fully extracted.

[0004] So far, researchers at home and abroad have proposed a series of methods for financial document-level event extraction. Although some progress has been made, the existing financial document-level event extraction methods still have many limitations, mainly concentrated in the following aspects:

[0005] 1. Dataset limitations: The existing financial document-level event extraction datasets are not perfect in terms of scale, fine-grained annotation, and event type coverage, making it difficult to fully support diverse event extraction needs.

[0006] 2. Insufficient ability to process noisy data: Financial documents usually contain a large amount of irrelevant information or noisy text, such as background descriptions or other non-event-related content. Existing models find it difficult to effectively filter out these noises, resulting in a decrease in the accuracy of event extraction.

[0007] 3. Limited understanding of complex financial terms: Financial documents often contain a large number of complex financial terms, such as financial indicators, professional terms, and industry-specific expressions. Existing models often have difficulty accurately interpreting these terms, resulting in poor performance of the model when processing texts involving professional knowledge.

[0008] 4. Weak ability to capture cross-sentence and cross-paragraph information: Financial events usually span multiple sentences or even paragraphs, and existing models cannot effectively capture the complete event semantics across sentences and paragraphs in long text analysis, resulting in one-sided extraction results and difficulty in fully reflecting the full picture of the event.

[0009] Therefore, finding a method that can both filter out noise in financial documents to improve the ability to understand financial documents and improve the accuracy of financial event extraction is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention

[0010] The present invention provides a financial document-level event extraction method and system based on multiple semantic enhancement, which is used to solve the defect of the prior art that the deep semantic information in the financial text cannot be captured, and to achieve the extraction of event information from multi-level and cross-paragraph texts when processing financial documents, thereby improving the accuracy and efficiency of event extraction.

[0011] The present invention provides a financial document-level event extraction method based on multiple semantic enhancement, comprising the following steps:

[0012] Constructing a financial document event model, wherein the financial document event model includes an event type set and an argument role set of the event type;

[0013] Acquire financial document data to be extracted, and preprocess and annotate the financial document data to be extracted based on event types and argument roles to obtain annotated financial document data;

[0014] The annotated financial document data is input into the trained financial document-level event extraction model, and the semantic matching optimization module, multi-granularity semantic enhancement module and event-centered heterogeneous graph structure module of the financial document-level event extraction model are used to extract financial document-level events from the annotated financial document data to obtain extraction results including event types and event arguments; wherein the semantic matching optimization module is used to filter and optimize the annotated financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event-centered heterogeneous graph structure module is used to capture the interactive relationship between nodes in the annotated financial document data based on the output of the multi-granularity semantic enhancement module.

[0015] According to a financial document-level event extraction method based on multiple semantic enhancement provided by the present invention, the processing steps of the semantic matching optimization module are:

[0016] Use the financial encoder to contextually encode the annotated financial document data and event templates to obtain high-dimensional vector representations of sentences and event templates;

[0017] Perform semantic similarity calculation on the high-dimensional vector representation of the sentence and the high-dimensional vector representation of the event template, and match it with a preset event template similarity threshold to obtain a matching result. Filter the annotated financial documents based on the matching result to obtain an optimized document.

[0018] According to a financial document-level event extraction method based on multi-semantic enhancement provided by the present invention, the processing steps of the multi-granularity semantic enhancement module are as follows:

[0019] Use the optimized document as the input text, and through the financial encoder, obtain sentence-level semantic features and paragraph-level semantic features;

[0020] Use a bidirectional gated recurrent unit to perform sequential modeling on the sentence-level semantic features and paragraph-level semantic features to obtain sequential semantic features; the formula is:

[0021] ;

[0022] ;

[0023] Where represents the sentence-level semantic features after Bi-GRU modeling, represents the paragraph-level semantic features after Bi-GRU modeling, represents the operation function of the bidirectional gated recurrent unit;

[0024] Perform global information modeling on the sequential semantic features through the self-attention mechanism to obtain sentence-level global features and paragraph-level global features; the formula is:

[0025] ;

[0026] ;

[0027] Where represents the sentence-level global features, represents the paragraph-level global features, represents the self-attention calculation function applied to the sentence-level global features, represents the self-attention calculation function applied to the paragraph-level global features;

[0028] Use the bidirectional interactive attention mechanism to fuse the sentence-level global features and paragraph-level global features to obtain multi-granularity semantic fusion features, and perform entity recognition on the multi-granularity semantic fusion features to obtain an entity set.

[0029] A financial document-level event extraction method based on multi-semantic enhancement provided by the present invention, wherein the bidirectional interactive attention mechanism is used to fuse the sentence-level global feature and the paragraph-level global feature, specifically including:

[0030] Normalize the sentence-level global feature and the paragraph-level global feature respectively to obtain the sentence-level normalized feature and the paragraph-level normalized feature;

[0031] Based on multi-head attention, perform information interaction on the sentence-level normalized feature and the paragraph-level normalized feature to obtain the interaction feature from the sentence level to the paragraph level and the interaction feature from the paragraph level to the sentence level;

[0032] Sum the interaction feature from the sentence level to the paragraph level and the interaction feature from the paragraph level to the sentence level with the sentence-level normalized feature and the paragraph-level normalized feature respectively by residual, and perform normalization processing on the result of the residual sum to generate the sentence-level secondary normalized feature and the paragraph-level secondary normalized feature;

[0033] Input the sentence-level secondary normalized feature and the paragraph-level secondary normalized feature into the feed-forward neural network respectively and perform residual connection and layer normalization operations to obtain the sentence-level interaction feature and the paragraph-level interaction feature;

[0034] Concatenate the sentence-level interaction feature and the paragraph-level interaction feature to obtain the multi-granularity semantic fusion feature;

[0035] Based on the multi-granularity semantic fusion feature, perform named entity recognition, and decode through the conditional random field to obtain the entity set.

[0036] A financial document-level event extraction method based on multi-semantic enhancement provided by the present invention, the processing steps of the event center heterogeneous graph structure module are:

[0037] Construct a heterogeneous graph structure including event nodes, entity nodes and sentence nodes based on the multi-granularity semantic fusion feature and the entity set, wherein the edges between nodes in the heterogeneous graph structure include the edges between events, the edges between events and entities, the edges between events and sentences, the edges between entities, the edges between sentences and the self-loop edges of events;

[0038] Use the graph neural network to perform feature propagation and update on the constructed heterogeneous graph structure to obtain the node representation integrating global relationship information.

[0039] A financial document-level event extraction method based on multi-semantic enhancement provided by the present invention, the training process of the financial document-level event extraction model includes:

[0040] Obtain financial document-level event data, preprocess and annotate the financial document-level event data to form the training set data of the financial document-level event extraction model, and use the event types and arguments annotated in each document of the training set data as the true event set;

[0041] Use the financial document-level event extraction model during training to extract events from the training set data to obtain a predicted event set;

[0042] Calculate the matching degree between the predicted event set and the true event set using the Hausdorff distance to obtain the Hausdorff distance loss;

[0043] Based on the Hausdorff distance and the entity representation loss, construct an overall loss function; among them, the entity representation loss is calculated based on the entity set extracted by the multi-granularity semantic enhancement module; the calculation formula of the overall loss function is:

[0044] ;

[0045] ;

[0046] Among them, \(L\) represents the overall loss function, represents the predicted event set and the true event set the Hausdorff distance loss between them, represents the predicted event set, represents the true event set, represents the entity representation loss, represents the matching pair of the predicted event and the true event, represents the set of all matching event pairs, represents the cross-entropy loss of the event type, represents the event type of the \(i\)-th predicted event, represents the event type of the \(j\)-th true event, represents the set of arguments in the event, \(e\) represents the argument, and \(k\) represents the index of the argument, represents the type of the \(k\)-th argument in the \(i\)-th predicted event, represents the type of the \(k\)-th argument in the \(j\)-th true event;

[0047] Optimize the overall loss function through the backpropagation algorithm, update the model parameters, and improve the accuracy of event extraction.

[0048] According to a financial document-level event extraction method based on multiple semantic enhancements provided by the present invention, the entity representation loss includes a sequence annotation loss function and a binary classification loss function; among them,

[0049] The sequence labeling loss function is determined based on named entity recognition of multi-granularity semantic features;

[0050] The binary classification loss function is determined using binary cross-entropy based on the probability that any two entities in the entity set belong to the same event.

[0051] The present invention also provides a financial document-level event extraction system based on multi-semantic enhancement, which implements the financial document-level event extraction method as described above, including:

[0052] An event pattern construction module for constructing a financial document event pattern, where the financial document event pattern includes an event type set and an argument role set of the event type;

[0053] A document processing module for preprocessing and annotating the financial document data to be extracted with event types and argument roles, and obtaining the annotated financial document data;

[0054] An event extraction module for inputting the annotated financial document data into a trained financial document-level event extraction model, and using the semantic matching optimization module, multi-granularity semantic enhancement module, and event center heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data, and obtaining an extraction result including event types and event arguments. The semantic matching optimization module is used to filter and optimize the annotated financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event center heterogeneous graph structure module is used to capture the interaction relationship between nodes in the annotated financial document data based on the output of the multi-granularity semantic enhancement module.

[0055] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the financial document-level event extraction method as described in any one of the above.

[0056] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the financial document-level event extraction method as described in any one of the above.

[0057] The financial document-level event extraction method and system based on multi-semantic enhancement provided by the present invention effectively handle the long-distance dependence and cross-paragraph information integration problems in document-level event extraction through the semantic matching optimization module, multi-granularity semantic enhancement module, and event center heterogeneous graph structure module, improve the extraction accuracy of complex financial events, and realize the efficient extraction of event information in financial documents. Description of the Drawings

[0058] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 is a flowchart of the financial document-level event extraction method provided by the present invention;

[0060] Figure 2 is a flowchart of the financial document-level event extraction method provided by the present invention;

[0061] Figure 3 is a flowchart of the process of fusing and generating multi-granularity semantic fusion features using a bidirectional interactive attention mechanism in the financial document-level event extraction method provided by the present invention;

[0062] Figure 4 is a schematic diagram of the training set data and annotation examples of the financial document-level event extraction method provided by the present invention;

[0063] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0064] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0065] As Figure 1 and Figure 2 shown, the present invention provides a financial document-level event extraction method based on multi-semantic enhancement, including the following steps:

[0066] S1. Construct a financial document event pattern, where the financial document event pattern includes an event type set and an argument role set of the event type;

[0067] S2. Obtain the financial document data to be extracted, preprocess and annotate the financial document data to be extracted based on the event type and argument role, and obtain the annotated financial document data;

[0068] S3. Input the labeled financial document data into the trained financial document-level event extraction model, and use the semantic matching optimization module, multi-granularity semantic enhancement module, and event-centered heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the labeled financial document data, obtaining an extraction result including event types and event arguments; wherein the semantic matching optimization module is used to filter and optimize the labeled financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event-centered heterogeneous graph structure module is used to capture the interaction relationships between nodes in the labeled financial document data based on the output of the multi-granularity semantic enhancement module.

[0069] The present invention effectively addresses the problems of long-distance dependencies and cross-paragraph information integration in document-level event extraction through the semantic matching optimization module, multi-granularity semantic enhancement module, and event-centered heterogeneous graph structure module, improves the extraction accuracy of complex financial events, and realizes the efficient extraction of event information in financial documents.

[0070] In an embodiment of the present invention, the financial document event pattern includes 22 types of financial event types and 116 argument roles, which can effectively summarize the actual business requirements of financial enterprises, making the financial document event pattern comprehensive and detailed.

[0071] In an embodiment of the present invention, the processing steps of the semantic matching optimization module are as follows:

[0072] Use a financial encoder to perform context encoding on the labeled financial document data and event templates, obtaining sentence high-dimensional vector representations and event template high-dimensional vector representations;

[0073] Calculate the semantic similarity between the sentence high-dimensional vector representations and the event template high-dimensional vector representations, and match the result with a preset event template similarity threshold to obtain a matching result. Filter the labeled financial document based on the matching result to obtain an optimized document.

[0074] In an embodiment of the present invention, the financial encoder can use Mengzi-FinBERT (Mengzi-BERT-base-fin). Mengzi-FinBERT is a financial-specific pre-trained model based on the RoBERTa (A Robustly Optimized BERT Pretraining Approach) architecture, which has been fine-tuned with a large amount of data in the financial field to enhance the model's ability to process financial text.

[0075] Specifically, a specific embodiment is used to illustrate the processing steps of the semantic matching optimization module:

[0076] Financial document , where represents the i-th sentence in the financial document, n represents the total number of sentences in the financial document, and the event template is , where represents the j-th event type, m represents the total number of event types in the event template, and each event type contains a specific set of argument roles ;

[0077] The financial encoder Mengzi-FinBERT is used to encode each sentence in the financial document and the event template word by word to obtain the high-dimensional vector representation of the sentence and the high-dimensional vector representation of the event template :

[0078]

[0079]

[0080] where and represent the high-dimensional vector representations of the sentence and the event template respectively, represents the high-dimensional vector representation of the sentence , represents the high-dimensional vector representation of the event template ;

[0081] The semantic similarity between each sentence and the event template is measured by the cosine similarity, and its similarity calculation formula is as follows:

[0082]

[0083] where, represents the semantic similarity between the sentence and the event template , represents the L2 norm;

[0084] According to the calculated similarity value, a similarity threshold is set, and the semantic similarity is compared with the similarity threshold:

[0085] When , the sentence is retained;

[0086] Otherwise, the sentence is filtered out ;

[0087] Based on the similarity threshold for screening, a set of sentences D' that highly matches the event template T is extracted from the financial document D, that is:

[0088]

[0089] The sentences in the filtered document D' are all content highly relevant to financial events, ensuring the effectiveness and pertinence of subsequent processing.

[0090] It can be understood that the similarity threshold is set according to actual usage requirements, and the present invention does not make specific limitations on this.

[0091] The semantic matching optimization module of the present invention uses Mengzi-FinBERT for context encoding, filters out irrelevant sentences by calculating the semantic similarity between sentences in the financial document and the event template, effectively improving the understanding ability of the financial document-level event extraction model for financial professional terms and reducing the interference of noise data on financial document-level event extraction.

[0092] In an embodiment of the present invention, the processing steps of the multi-granularity semantic enhancement module are as follows:

[0093] Taking the optimized document as the input text, through the financial encoder, sentence-level semantic features and paragraph-level semantic features are obtained;

[0094] Using a bidirectional gated recurrent unit to perform sequential modeling on the sentence-level semantic features and paragraph-level semantic features to obtain sequential semantic features;

[0095] Performing global information modeling on the sequential semantic features through a self-attention mechanism to obtain sentence-level global features and paragraph-level global features;

[0096] Using a bidirectional interactive attention mechanism to fuse the sentence-level global features and paragraph-level global features to obtain multi-granularity semantic fusion features, and performing entity recognition on the multi-granularity semantic fusion features to obtain an entity set.

[0097] It can be understood that using a bidirectional gated recurrent unit for sequential modeling can capture the context dependence in the features and integrate context information.

[0098] Specifically, a specific embodiment is used to illustrate the processing steps of the multi-granularity semantic enhancement module:

[0099] Taking the optimized document as the input text, through the financial encoder, sentence-level semantic features and paragraph-level semantic features ;

[0100] Serial modeling is performed using a bidirectional gated recurrent unit (Bi-GRU, Bidirectional Gated Recurrent Unit), and the formula is:

[0101] ;

[0102] ;

[0103] Among them, represents the sentence-level semantic features after Bi-GRU modeling, represents the paragraph-level semantic features after Bi-GRU modeling, represents the operation function of the bidirectional gated recurrent unit;

[0104] The self-attention mechanism is used to capture the correlation between each feature vector and other feature vectors, and the formula is:

[0105] ;

[0106] ;

[0107] Among them, represents the sentence-level global features, represents the paragraph-level global features, represents the self-attention calculation function applied to the sentence-level global features, represents the self-attention calculation function applied to the paragraph-level global features.

[0108] Furthermore, the bidirectional interactive attention mechanism is used to fuse the sentence-level global features and the paragraph-level global features, which specifically includes:

[0109] Normalize the sentence-level global features and the paragraph-level global features respectively to obtain sentence-level normalized features and paragraph-level normalized features;

[0110] Based on multi-head attention, perform information interaction on the sentence-level normalized features and the paragraph-level normalized features to obtain interaction features from the sentence level to the paragraph level and interaction features from the paragraph level to the sentence level;

[0111] Sum the interaction features from the sentence level to the paragraph level and the interaction features from the paragraph level to the sentence level with the sentence-level normalized features and the paragraph-level normalized features respectively by residual, and normalize the results of the residual sum to generate sentence-level secondary normalized features and paragraph-level secondary normalized features;

[0112] Input the sentence-level second normalization feature and the paragraph-level second normalization feature into the feed-forward neural network respectively, and perform residual connection and layer normalization operations to obtain the sentence-level interaction feature and the paragraph-level interaction feature;

[0113] Concatenate the sentence-level interaction feature and the paragraph-level interaction feature to obtain the multi-granularity semantic fusion feature;

[0114] Based on the multi-granularity semantic fusion feature, perform named entity recognition and decode through the conditional random field to obtain the entity set.

[0115] Specifically, a specific embodiment is used for illustration:

[0116] Such as Figure 3 shown, process the sentence-level feature and the paragraph-level feature through layer normalization (LayerNormalization, LN) respectively to obtain the sentence-level normalized feature and the paragraph-level normalized feature. The calculation formula is:

[0117]

[0118]

[0119] Among them, represents the sentence-level normalized feature, represents the paragraph-level normalized feature, represents the sentence-level feature, represents the paragraph-level feature, represents the layer normalization function;

[0120] Input the sentence-level normalized feature and the paragraph-level normalized feature into the multi-head attention (Multi-Head Attention, MHA), and use the sentence-level feature as the query matrix , and use the paragraph-level feature as the key matrix and the value matrix . The interaction feature from the paragraph level to the sentence level is:

[0121]

[0122] Among them, represents the interaction feature from the paragraph level to the sentence level, represents the multi-head attention mechanism, represents the concatenation operation, represents the output feature of the i-th attention head, represents the output feature of the h-th attention head, and h represents the total number of attention heads. A weight matrix for projecting the concatenated multi-head attention output into the target dimensional space;

[0123] Regarding the paragraph-level features as the query matrix and the sentence-level features as the key matrix and the value matrix respectively, the interactive features from the sentence level to the paragraph level are:

[0124]

[0125] where represents the interactive features from the sentence level to the paragraph level;

[0126] where the specific calculation for each attention head is:

[0127]

[0128] where represents the softmax function operation of the activation function, represents the query matrix of the i-th attention head, represents the transpose of the key matrix of the i-th attention head, represents the dimension of the key vector, represents the value matrix of the i-th attention head;

[0129] The interactive features from the sentence level to the paragraph level and the interactive features from the paragraph level to the sentence level are respectively summed with the sentence-level normalized features and the paragraph-level normalized features by residual, and the result of the residual sum is normalized to generate the sentence-level second-normalized features and the paragraph-level second-normalized features. The calculation formula is:

[0130]

[0131]

[0132] where represents the sentence-level second-normalized features, represents the paragraph-level second-normalized features;

[0133] The sentence-level second-normalized features and the paragraph-level second-normalized features are respectively input into the Feedforward Neural Network (FNN) and subjected to residual connection and layer normalization operations to obtain the sentence-level interactive features and the paragraph-level interactive features. The calculation formula is:

[0134]

[0135]

[0136] Among them, represents the sentence-level interaction feature, represents the paragraph-level interaction feature, represents the feed-forward neural network operation;

[0137] The sentence-level interaction feature and the paragraph-level interaction feature are concatenated to generate a multi-granularity semantic fusion feature , and the calculation formula is:

[0138]

[0139] In an embodiment of the present invention, the BIO annotation system is used to perform named entity recognition on the multi-granularity semantic fusion feature , decoded by a conditional random field, and the entity set in the financial document is extracted. The calculation formula is:

[0140]

[0141] Among them, represents the entity set, represents the decoding operation of the conditional random field.

[0142] The multi-granularity semantic enhancement module of the present invention captures sentence-level and paragraph-level semantic features through BiGRU and self-attention mechanism respectively, and uses bidirectional interactive attention mechanism for feature fusion, realizing the effective integration of fine-grained and coarse-grained semantic information, and enhancing the capture ability of the financial document-level event extraction model for cross-sentence and cross-paragraph event information.

[0143] In an embodiment of the present invention, the processing steps of the event center heterogeneous graph structure module are as follows:

[0144] Based on the multi-granularity semantic fusion feature and the entity set, a heterogeneous graph structure including event nodes, entity nodes and sentence nodes is constructed. Among them, the edges between nodes in the heterogeneous graph structure include the edges between events, the edges between events and entities, the edges between events and sentences, the edges between entities, the edges between sentences, and the self-loop edges of events;

[0145] The constructed heterogeneous graph structure is propagated and updated with graph neural network to obtain the node representation integrating global relationship information.

[0146] It is understandable that the present invention integrates various information types in financial documents into interrelated nodes and edges, constructing an event-centered heterogeneous graph structure. Through the heterogeneous graph structure, the relationships between different types of nodes (such as entities, sentences, events) can be fully utilized to model complex semantics more comprehensively.

[0147] Specifically, the processing steps of the event-centered heterogeneous graph structure module are described with a specific embodiment:

[0148] Define the heterogeneous graph structure , where V is the set of nodes, including event nodes , entity nodes and sentence nodes , where the event nodes represent the semantic core of the event, randomly generate the initial embedding of the event nodes , and learn reasonable event semantic representations through training to capture the event context; entity nodes are extracted through named entity recognition in the multi-granularity semantic enhancement module, and the initial embedding of the entity nodes is determined based on the entity set , which is used to transfer entity information to the event nodes; sentence nodes represent document sentences, and the initial embedding of the sentence nodes is determined based on the sentence-level features , providing context semantic information for the event nodes, is the set of edges, including the bidirectional edges of the event nodes , event self-loop edges , directed edges from entities to event nodes , directed edges from sentences to event nodes , bidirectional edges between entity nodes and bidirectional edges between sentence nodes ; among them, is used for information exchange between event nodes; connects the event node itself, used to capture the features and information of the event node itself, is used to transfer entity information to the event node; is used to provide sentence context information to the event node, is used to model the semantic association between entities and strengthen the information sharing between entities; is used to model the context semantic association between sentences and enhance the context information transfer between sentences;

[0149] Perform event decoding on the event node representation, dividing the event nodes into two parts: event type classification and event argument classification. Among them, the event type classification is obtained by taking the final representation vector of the i-th event node Input a multi-layer perceptron (MLP) and use the softmax function to obtain the probability distribution of event types ; where is the initial embedding obtained after graph neural network message passing and update; the calculation formula of the probability distribution is:

[0150]

[0151] where, represents the probability distribution of event types, represents the softmax function, represents the multi-layer perceptron, represents the final representation vector of the i-th event node;

[0152] Event argument classification aggregates multiple mentions of entities through the multi-head attention mechanism (MHA), with the event node as the query, and the entity mentions as the key and value, to obtain the aggregated representation of the entity:

[0153]

[0154] where, represents the entity under the event node the semantic vector representation, represents the multi-head attention mechanism, represents taking the event node as the query (Query) representation, represents taking the entity as the key (Key) representation, represents taking the entity as the value (Value) representation;

[0155] Combining the aggregated representations of the event node and the entity through the MLP, calculate the type probability distribution of each argument :

[0156]

[0157] where, represents the entity under the event node the aggregated representation;

[0158] Determine the event type and the argument , the formula is:

[0159]

[0160]

[0161] Among them, represents the predicted event type, represents the predicted argument type, represents taking the maximum value, represents the type probability distribution of the k-th argument of the i-th event;

[0162] It can be understood that the event-centered heterogeneous graph structure module supports the interactive modeling of complex events in financial documents and improves the event extraction effect. The update of node features is achieved through a graph neural network (GNN-FiLM, Graph Neural Network with Feature-wise Linear Modulation), and the feature linear modulation mechanism (FiLM) is used to modulate the node features. The feature representation of each node is updated through message aggregation of neighboring nodes, and the feature linear modulation mechanism further modulates the node features, finely adjusts the node features through scaling and offset parameters, integrates the semantic information from entity nodes and context nodes, and generates an event-related representation.

[0163] The event heterogeneous graph of the present invention includes three types of nodes: entities, sentences, and events, as well as various types of edge relationships. The node features are updated through a graph neural network and a feature linear modulation mechanism, effectively modeling the complex interaction relationships between various types of information in the document and improving the recognition accuracy of event elements.

[0164] In an embodiment of the present invention, the training process of the financial document-level event extraction model includes:

[0165] Obtain financial document-level event data, preprocess and annotate the financial document-level event data to form the training set data of the financial document-level event extraction model, and use the event type and arguments annotated in each document of the training set data as the true event set;

[0166] Use the financial document-level event extraction model during training to extract events from the training set data to obtain a predicted event set;

[0167] Calculate the matching degree between the predicted event set and the true event set using the Hausdorff distance to obtain the Hausdorff distance loss;

[0168] Based on the Hausdorff distance and the entity representation loss, construct an overall loss function; among them, the entity representation loss is calculated based on the entity set extracted by the multi-granularity semantic enhancement module; the calculation formula of the overall loss function is:

[0169] ;

[0170] ;

[0171] Among them, L represents the overall loss function, represents the predicted event set and the true event set the Hausdorff distance loss between them, represents the predicted event set, represents the true event set, represents the entity representation loss, represents the matching pairs of predicted events and true events, represents the set of all matching event pairs, represents the cross-entropy loss of event types, represents the event type of the i-th predicted event, represents the event type of the j-th true event, represents the set of arguments in the event, e represents the argument, and k represents the index of the argument, represents the type of the k-th argument in the i-th predicted event, represents the type of the k-th argument in the j-th true event;

[0172] Optimize the overall loss function through the backpropagation algorithm, update the model parameters, and improve the accuracy of event extraction.

[0173] As Figure 4 shown, specifically, a specific embodiment is used for illustration:

[0174] According to the business requirements of financial enterprises, financial events are divided into six categories, including equity changes, company changes, performance changes, stock market fluctuations, executive changes, business changes, and regulatory penalties. At the same time, the concept of opposing events is introduced, such as losses and profits, stock price increases and decreases, etc. By defining 22 event types and 116 argument roles, it is ensured that the event types and argument roles do not overlap and are clear, thus constructing a comprehensive event pattern to meet the complex scenario requirements in the financial field; among them, the 22 event types are divided into 6 major categories;

[0175] Collect data from seven financial news websites such as Sina Finance, Flush Finance, and China Economic Net;

[0176] First, exclude duplicate news reports through deduplication; then, conduct format checks and delete documents with non-standard formats (such as articles in PDF or table form); next, filter out irrelevant content and exclude non-financial advertisements and personal opinion articles; according to the exchange announcements, screen out data of companies involved in financial fraud issues to avoid false information interfering with model training; finally, conduct manual review to ensure the relevance, accuracy, and authenticity of the data;

[0177] Use the doccano annotation tool to annotate the collected data. The annotation process is divided into two main stages: event classification and argument extraction. In the event classification stage, assign corresponding event labels to each document according to predefined event types and mark the trigger words. In the argument extraction stage, annotate the argument entities of each event according to the argument roles defined in the event pattern. The annotation process includes an expert review and cross-validation mechanism to ensure the accuracy and consistency of the annotated data and generate high-quality financial document event annotation data.

[0178] Use the annotated data to train the financial document-level event extraction model to obtain a trained financial document-level event extraction model.

[0179] Furthermore, the entity representation loss includes a sequence annotation loss function and a binary classification loss function. Among them,

[0180] The sequence annotation loss function is determined based on named entity recognition of multi-granularity semantic features.

[0181] The binary classification loss function is determined using binary cross-entropy based on the probability that any two entities in the entity set belong to the same event.

[0182] In an embodiment of the present invention, the calculation formula based on the probability that any two entities in the entity set belong to the same event is as follows:

[0183]

[0184] Among them, represents the entity and the probability that entity j belongs to the same event. is the Sigmoid function. represents the concatenation operation. represents the representation vector of entity i. represents the representation vector of entity j.

[0185] The binary classification loss function uses binary cross-entropy (CE), and the specific calculation formula is:

[0186]

[0187] Among them, represents the binary classification loss function. represents the event association label between entity i and entity j. represents the binary cross-entropy loss function.

[0188] The total loss function of the entity representation is:

[0189]

[0190] Among them, represents the total loss function of entity representation, represents the sequence labeling loss function, represents the binary classification loss function.

[0191] It can be understood that the sequence labeling loss function is generated by using the BIO labeling system to identify named entities for multi-granularity semantic features.

[0192] The present invention provides a financial document-level event extraction system based on multi-semantic enhancement to implement the above-mentioned financial document-level event extraction method, including:

[0193] An event pattern construction module for constructing a financial document event pattern, where the financial document event pattern includes an event type set and an argument role set of the event type;

[0194] A document processing module for preprocessing and annotating the financial document data to be extracted with event types and argument roles to obtain the annotated financial document data;

[0195] An event extraction module for inputting the annotated financial document data into a trained financial document-level event extraction model, and using the semantic matching optimization module, multi-granularity semantic enhancement module, and event center heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data to obtain an extraction result including event types and event arguments; wherein the semantic matching optimization module is used to filter and optimize the annotated financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event center heterogeneous graph structure module is used to capture the interaction relationship between nodes in the annotated financial document data based on the output of the multi-granularity semantic enhancement module.

[0196] The following describes the financial document-level event extraction device provided by the present invention. The financial document-level event extraction device described below can be correspondingly referred to the financial document-level event extraction method described above.

[0197] Figure 5 Illustrates a schematic diagram of the entity structure of an electronic device, such as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the financial document-level event extraction method, which includes: constructing a financial document event pattern, obtaining the financial document data to be extracted, preprocessing and annotating the financial document data to be extracted based on the event type and argument role to obtain the annotated financial document data, inputting the annotated financial document data into the trained financial document-level event extraction model, and using the semantic matching optimization module, the multi-granularity semantic enhancement module, and the event center heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data to obtain an extraction result including the event type and event arguments.

[0198] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0199] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the financial document-level event extraction method provided by each of the above methods. The method includes: constructing a financial document event pattern, obtaining financial document data to be extracted, preprocessing and annotating the financial document data to be extracted based on event types and argument roles to obtain annotated financial document data, inputting the annotated financial document data into a trained financial document-level event extraction model, and using the semantic matching optimization module, multi-granularity semantic enhancement module, and event center heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data to obtain an extraction result including event types and event arguments.

[0200] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the financial document-level event extraction method provided by each of the above methods. The method includes: constructing a financial document event pattern, obtaining financial document data to be extracted, preprocessing and annotating the financial document data to be extracted based on event types and argument roles to obtain annotated financial document data, inputting the annotated financial document data into a trained financial document-level event extraction model, and using the semantic matching optimization module, multi-granularity semantic enhancement module, and event center heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data to obtain an extraction result including event types and event arguments.

[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A financial document-level event extraction method based on multiple semantic enhancement, characterized in that: The following steps are involved: Constructing a financial document event model, wherein the financial document event model includes an event type set and an argument role set of the event type; Acquire financial document data to be extracted, and preprocess and annotate the financial document data to be extracted based on event types and argument roles to obtain annotated financial document data; Input the annotated financial document data into the trained financial document-level event extraction model, and use the semantic matching optimization module, multi-granularity semantic enhancement module and event-centered heterogeneous graph structure module of the financial document-level event extraction model to extract financial document-level events from the annotated financial document data, and obtain extraction results including event types and event arguments; wherein the semantic matching optimization module is used to filter and optimize the annotated financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event-centered heterogeneous graph structure module is used to capture the interactive relationship between nodes in the annotated financial document data based on the output of the multi-granularity semantic enhancement module; the processing steps of the multi-granularity semantic enhancement module are: using the optimized document as input text, passing it through a financial encoder, and obtaining sentence-level semantic features and paragraph-level semantic features; The sentence-level semantic features and paragraph-level semantic features are serialized and modeled using a bidirectional gated recurrent unit to obtain serialized semantic features; the formula is: ; ; in, Represents the sentence-level semantic features after Bi-GRU modeling, represents the paragraph-level semantic features after Bi-GRU modeling, represents the operation function of a bidirectional gated recurrent unit; The global information of the serialized semantic features is modeled through the self-attention mechanism to obtain sentence-level global features and paragraph-level global features; the formula is: ; ; in, represents the sentence-level global feature, represents the paragraph-level global features, represents the self-attention calculation function applied to sentence-level global features, represents the self-attention calculation function applied to paragraph-level global features; The sentence-level global features and paragraph-level global features are fused using a bidirectional interactive attention mechanism to obtain multi-granularity semantic fusion features, and entity recognition is performed on the multi-granularity semantic fusion features to obtain an entity set.

2. The financial document-level event extraction method based on multiple semantic enhancement according to claim 1 is characterized in that: The processing steps of the semantic matching optimization module are: Use the financial encoder to contextually encode the annotated financial document data and event templates to obtain high-dimensional vector representations of sentences and event templates; The semantic similarity between the sentence high-dimensional vector representation and the event template high-dimensional vector representation is calculated, and matched with a preset event template similarity threshold to obtain a matching result, and the annotated financial documents are filtered based on the matching result to obtain an optimized document.

3. The financial document-level event extraction method based on multiple semantic enhancement according to claim 1 is characterized in that: The method of fusing the sentence-level global features and the paragraph-level global features by using a bidirectional interactive attention mechanism specifically includes: Normalizing the sentence-level global features and the paragraph-level global features respectively to obtain sentence-level normalized features and paragraph-level normalized features; Based on multi-head attention, information interaction is performed on the sentence-level normalized features and the paragraph-level normalized features to obtain interactive features from the sentence level to the paragraph level and interactive features from the paragraph level to the sentence level; The interactive features from sentence level to paragraph level and the interactive features from paragraph level to sentence level are respectively summed with the sentence level normalized features and the paragraph level normalized features, and the results of the residual summation are normalized to generate sentence level secondary normalized features and paragraph level secondary normalized features; The sentence-level quadratic normalized features and the paragraph-level quadratic normalized features are respectively input into the feedforward neural network and subjected to residual connection and layer normalization operations to obtain sentence-level interaction features and paragraph-level interaction features; The sentence-level interaction features and the paragraph-level interaction features are concatenated to obtain multi-granularity semantic fusion features; Named entity recognition is performed based on the multi-granularity semantic fusion features, and an entity set is obtained through conditional random field decoding.

4. The financial document-level event extraction method based on multiple semantic enhancement according to claim 3 is characterized in that: The processing steps of the event center heterogeneous graph structure module are: Based on multi-granularity semantic fusion features and entity sets, a heterogeneous graph structure including event nodes, entity nodes and sentence nodes is constructed, wherein the edges between nodes in the heterogeneous graph structure include edges between events, edges between events and entities, edges between events and sentences, edges between entities and entities, edges between sentences and sentences, and self-loop edges of events; A graph neural network is used to propagate and update the features of the constructed heterogeneous graph structure to obtain a node representation that integrates global relationship information.

5. The financial document-level event extraction method based on multiple semantic enhancement according to claim 4 is characterized in that: The training process of the financial document-level event extraction model includes: Obtain financial document-level event data, preprocess and annotate the financial document-level event data, and form the training set data for the financial document-level event extraction model. The event type and argument annotated in each document of the training set data are used as the real event set. Use the financial document-level event extraction model in training to extract events from the training set data to obtain the predicted event set; The Hausdorff distance is used to calculate the matching degree between the predicted event set and the real event set, and the Hausdorff distance loss is obtained; Based on the Hausdorff distance and entity representation loss, an overall loss function is constructed. The entity representation loss is calculated based on the entity set extracted by the multi-granularity semantic enhancement module. The calculation formula of the overall loss function is: ; ; Among them, L represents the overall loss function, Represents the predicted event set With true events The Hausdorff distance loss between represents the set of predicted events, represents a set of real events, Represents an entity that represents a loss, represents the matching pair of predicted events and real events, Represents the set of all matching event pairs, represents the cross entropy loss of event types, represents the event type of the i-th predicted event, represents the event type of the jth real event, Represents the set of arguments in the event, e represents the argument, k represents the index of the argument, represents the type of the kth argument in the i-th predicted event, represents the type of the kth argument in the jth real event; The overall loss function is optimized through the back-propagation algorithm, the model parameters are updated, and the accuracy of event extraction is improved.

6. The financial document-level event extraction method based on multiple semantic enhancement according to claim 5 is characterized in that: The entity representation loss includes a sequence labeling loss function and a binary classification loss function; wherein, The sequence labeling loss function is determined based on named entity recognition of multi-granularity semantic features; The binary classification loss function is determined based on the probability that any two entities in the entity set belong to the same event using binary cross entropy.

7. A financial document-level event extraction system based on multiple semantic enhancement, characterized in that: Implementing the financial document-level event extraction method as described in any one of claims 1 to 6, comprising: An event pattern building module, used to build a financial document event pattern, wherein the financial document event pattern includes an event type set and an argument role set of the event type; A document processing module, used for preprocessing and labeling the financial document data to be extracted according to event type and argument role, to obtain labeled financial document data; An event extraction module is used to input the annotated financial document data into a trained financial document-level event extraction model, and use the semantic matching optimization module, multi-granularity semantic enhancement module and event-centered heterogeneous graph structure module of the financial document-level event extraction model to perform financial document-level event extraction on the annotated financial document data to obtain extraction results including event types and event arguments; wherein the semantic matching optimization module is used to filter and optimize the annotated financial document data, the multi-granularity semantic enhancement module is used to perform context encoding based on the output of the semantic matching optimization module to fuse sentence-level semantic features and paragraph-level semantic features and perform entity recognition, and the event-centered heterogeneous graph structure module is used to capture the interactive relationship between nodes in the annotated financial document data based on the output of the multi-granularity semantic enhancement module.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the financial document-level event extraction method as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the financial document-level event extraction method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and device for extracting chapter-level event based on multi-granularity entity heterogeneous graph

    CN114742016A

  • Event extraction method and device, electronic equipment and storage medium

    CN117493500A