Long text financial event extraction and key event extraction method and system based on small language model

By building an event interaction layer in financial announcements and using a small language model to complete information, the problem of insufficient information completion and event correlation analysis of event extraction in long-text financial announcements is solved, and the integrity of event description and financial analysis support is achieved.

CN119938899AActive Publication Date: 2025-05-06NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510038762.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The prior art has problems with insufficient information completion and event correlation analysis in event extraction in long-text financial announcements, especially in cross-document scenarios, it is difficult to accurately complete information and identify key events.

Method used

By building an event interaction layer, using a small language model to flow and complement each other between different documents, the integrity of event description is achieved, and a cross-document event interaction diagram is designed to identify key events.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of event extraction, ensures the completeness and coherence of event description, and provides financial analysis with richer event information support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938899A_ABST
    Figure CN119938899A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and discloses a long text financial event extraction and key event extraction method and system based on a small language model. According to the event extraction method for the long text financial announcement, the announcement content is extracted through the text abstract module, redundant information is removed, processing resources are reduced, and the extraction accuracy is improved. Meanwhile, aiming at the co-reference phenomenon of the small language model in single document event extraction, a co-reference ablation strategy is provided, so that the accuracy of event identification is effectively improved, and comparison experiments prove that the method provided by the invention improves the accuracy and recall rate of event element extraction. By constructing the event interaction layer, the flow and complementation of event information among multiple documents are realized, the integrity and continuity of event description are ensured, the comprehensiveness and accuracy of key event extraction are remarkably enhanced, and richer event information support is provided for financial analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for extracting long-text financial events and key events based on a small language model. Background Art

[0002] With the rapid evolution of financial markets, the amount of information contained in financial announcements continues to increase. These announcements often contain complex events, such as company mergers, shareholder changes, asset restructuring, etc., which are of great reference value to investors, regulators and market participants. The text of financial announcements is long and information-intensive. The content of events is presented in a variety of forms, including a large number of potentially influential market events. How to accurately extract key information from these long texts and perform structured processing to more efficiently support decision-making and analysis has become a key demand in the financial field.

[0003] At present, traditional event extraction methods are mostly based on rules and templates, mainly capturing specific types of events and their elements, but these methods have limitations when facing a rapidly changing market environment and diverse event content. In recent years, deep learning has promoted the development of natural language processing (NLP) technology, especially the application of small language models in text extraction, which has brought new ideas to event extraction. Compared with large models, these lightweight models have lower computational costs and smaller parameter requirements, and can process long text data more efficiently while ensuring a certain extraction effect. However, in long-text financial announcements, event extraction still faces significant challenges: existing methods are insufficient in information completion and event association analysis. Although some information extraction frameworks (such as UIE) support multi-event extraction, they still lack the ability to automatically complete missing information and in-depth analysis of the association between events in financial applications. Especially in cross-document scenarios, how to complete information and identify key events is still a key direction to improve the effect of event extraction. Summary of the invention

[0004] In view of the shortcomings of the above-mentioned prior art, the present invention proposes a method and system for extracting long text financial events and key events based on a small language model. By building an event interaction layer, the present invention can effectively flow information and complement each other between different documents, ensure the integrity of event descriptions, more effectively meet the market's demand for accurate and comprehensive event information, and promote the progress and development of financial analysis.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows: a method for extracting long text financial events and key events based on a small language model, comprising the following steps:

[0006] Step 1: Obtain and preprocess the financial announcement dataset to generate standardized announcement titles and announcement contents;

[0007] Step 2: Perform sentence embedding and summary generation on the data obtained in step 1 to obtain high-importance sentences to form a complete text summary S summary ;

[0008] Step 3: Summarize the text S summary Multi-label classification is performed and classified into the 13 financial event types set C = {c1, c2, ..., c 13}, including company listing, shareholder reduction, shareholder increase, enterprise acquisition, enterprise financing, share repurchase, pledge, contact pledge, enterprise bankruptcy, loss, interview, winning bid, and senior management changes;

[0009] Step 4: Fine-tune the small language model based on the UIE general information extraction framework, extract financial events in detail, and merge co-reference events in the same document through the event co-reference resolution method; the focus of the fine-tuned small language model is to be responsible for converting text information into structured extracted information.

[0010] Step 5: Design an event interaction layer for cross-document event completion and key event identification to obtain the missing information completion and top-k key event information of the same listed company across documents.

[0011] The step 1 specifically includes:

[0012] Step 1.1: Data acquisition and format conversion: crawl the financial announcement data of listed companies through the data acquisition module, extract key information, and save it as a csv file;

[0013] Step 1.2: Trace the announcement title in the csv file title And the announcement content T content The fields are cleaned to remove non-structured symbols, HTML tags and irrelevant characters in the fields. The cleaned announcement title and announcement content are recorded as and

[0014] The step 2 specifically includes:

[0015] Step 2.1: Sentence segmentation and embedding generation: Divide into sentence set T = {s1, s2, ..., sn}, use the pre-trained model bert-base-chinese to generate sentence embedding; for each sentence s i , use bert-base-chinese to generate a d-dimensional vector representation h i , we get the implicit representation matrix H = {h1, h2, ..., hn} of the text, where each Representative sentence s i Embedding

[0016] Step 2.2: Sentence importance calculation: Calculate each sentence s i The importance weight w(s i ) to select representative sentences; use TF-IDF weight w(s i ) to calculate: Among them, TF(w) represents the word frequency of word w, and IDF(w) is the inverse document frequency;

[0017] Step 2.3: Summary generation: According to the weight w(s) of each sentence i ) are sorted, and the top m important sentences are selected from high to low to form a summary candidate set And keep the original order of the sentence, and convert the candidate set S candidate The sentences in are concatenated in order to form a complete summary:

[0018] The step 3 specifically includes:

[0019] Step 3.1: Feature vectorization: For the complete summary S summary Use the bert-base-chinese model for encoding and embed the text into a feature vector representation F = {f1, f2, ..., f d}, d is the feature dimension, and the feature vector represents the semantic information that can preserve the event category;

[0020] Step 3.2: Multi-label classification layer setting: Introduce the multi-label classification head of the bert-base-chinese model and convert the output feature vector F = {f1, f2, ..., f d}, input to the multi-label classification head to generate the category classification score The specific calculation formula is: The weight matrix and bias term are automatically adjusted during the bert-base-chinese model training process, and the binary cross entropy loss function is used to optimize the multi-label classification head. For each category, the prediction probability p(c i |T) and the true label y i , the BCE loss function is calculated as: y i =1 when the text belongs to category c1;

[0021] Step 3.3: Normalization and threshold determination: Use the sigmoid function to map the category classification scores to probabilities Calculate the probability of each category independently and set a threshold τ to determine whether the text belongs to a specific category c i, the final multi-label classification result is expressed as:

[0022] The step 4 specifically includes:

[0023] Step 4.1: Data annotation: Generate 10-shot samples for 13 types of financial events, annotate them on the doccano data annotation platform, and store the annotation files in JSON format, including event categories, trigger words, arguments, and argument positions;

[0024] Step 4.2: Fine-tune the UIE-Base model as a financial event extraction model: Use the UIE module in the Python PaddleNLP library for supervised full-scale fine-tuning, use AutoTokenizer to load the tokenizer from the pre-trained UIE-Bas model, load_dataset to load the training dataset and validation dataset, and use the convert_example function to convert them as model input; set the hyperparameters at the same time;

[0025] Step 4.3: Financial event extraction and similarity calculation: According to the event classification results, they are divided into 13 categories of extraction templates. The Taskflow library under the PaddleNLP framework is applied to call the financial event extraction model fine-tuned in step 4.2; in the extraction task, Taskfiow is called, the parameter is set to information_extraction, and the schema is put into the financial event extraction model in the form of prompt to start extraction; in the first round of extraction results, similarity calculations are performed on similar events extracted from the same text to determine whether they are the same event; first, the pre-trained model all-MiniLM-L6-v2 in the Sentence Transformers library is used to encode the text arguments {t1, t2, t3} into vector representations {v1, v2, v3}, v i =Encode(t1); use cosine similarity to calculate the similarity between encoded vectors

[0026] Step 4.4: Co-reference resolution and output formatting: For similar events, first set a threshold threshold to control the merging condition; if the similarity is higher than the threshold, it is considered as the same event chain, and output a list L containing similar event pairs and their similarities = {(e i , e j ,cosine_similarity(e i , e j ))}Select the event with the best expression completeness as the benchmark template base; The completeness of the event argument set R(e) = {r1, r2, ..., r n}, where r i is an argument; the argument r is missing in the benchmark event j , then complete it from other events and select the argument with the largest probability value as the final result, which is expressed as: ), where P(r j |e i ) indicates that in event e i Zhonglun Yuan j The merged events will be stored as DataFrame together with the announcement type, announcement title, announcement content, company name and event time.

[0027] {Announcement type, announcement title, announcement content, company name, event time, r1, r2, ...}.

[0028] The step 5 specifically includes:

[0029] Step 5.1: Cross-document event similarity calculation and completion: Each event is regarded as a node e i , use SentenceTransformers to transform e i Coded as V i ; Use Euclidean distance to calculate the distance between any two nodes e i and e j The similarity S ij , d ij =||V i -V j ||,S ij The larger the event e i and e j The more similar they are; all event nodes and their similarity edge weights are constructed into an event interaction graph G = (E, S), where E is the event node set and S is the edge weight matrix;

[0030] Step 5.2: Key event screening and completion: For each event node e i Calculate its relevance score R(e i )=∑ j∈N(i) S ij , where N(i) represents the node e i The set of adjacent nodes; the higher the correlation score, the stronger the correlation of the event in the network; according to the correlation score R(e i ), select the first k event nodes to form a key event chain, denoted as set ε key ; For each node e in the key event chain k ∈εkey , whose arguments are the set Rk={r k,1 , r k,2 , ..., r k,m}; Event node e k When an argument is missing from , find the missing argument: Among them, P(r j ) is the argument r j The probability of occurrence in the adjacent nodes, N(k) represents the set of adjacent nodes with high similarity; according to the edge weight S kj The size of the argument is used to select the most reliable candidate completion argument; the argument value r with the highest edge weight among the adjacent nodes is selected j,m+1 To complete e k After completing all key event nodes, the complete event chain information is output.

[0031] A system for extracting financial events and key events from long texts based on a small language model, which uses a method for extracting financial events and key events from long texts based on a small language model; the system for extracting financial events and key events from long texts based on a small language model includes a data cleaning module, a long text summary module, an event classification module, an event extraction and coreference resolution module, a cross-document event interaction layer, and a key event extraction module;

[0032] S1. Obtain the financial announcement data of listed companies through the data cleaning module and perform cleaning processing to generate standardized announcement titles and announcement contents;

[0033] S2. Input the cleaned and standardized announcement title and content into the long text summary module to generate sentence vector representation and obtain high-importance sentences to form a complete summary;

[0034] S3, input the complete summary content into the event classification module for encoding, generate feature vectors, optimize model performance through the classification layer, and obtain multi-category event classification results;

[0035] S4. Input the financial announcement summary and event classification results into the event extraction and coreference resolution module, apply the small language model after annotation and fine-tuning through the UIE framework to perform event extraction and coreference event resolution, obtain complete structured event information and save it;

[0036] S5. Input the structured event information into the cross-document event interaction layer and the key event extraction module to obtain the missing information completion and top-k key event information of the same listed company across documents, and store the processed structured event information into the database.

[0037] The data cleaning module is responsible for obtaining the csv format data of the listed company's financial announcements, extracting key information, and saving it as a standard format file; the data cleaning module cleans the announcement title and content, removes non-structured symbols, tags and irrelevant characters, generates cleaned announcement title and content, and provides a standardized data basis for subsequent processing;

[0038] The long text summary module includes a text embedding and importance calculation unit and a long text summary module; the text embedding and importance calculation unit is responsible for segmenting the cleaned announcement content into sentences and generating vectorized representations of the sentences; the long text summary module further determines the relative importance of each sentence based on a weight calculation method; the summary generation unit selects sentences with high document frequency, sorts them by importance, and finally generates a complete summary to highlight the core content of the text;

[0039] The event classification module is responsible for encoding the summary content and converting the text into a feature vector; the event classification module sets a classification layer so that the model generates classification scores for various types of events and optimizes model performance; through normalization operations, the event classification module converts the classification scores into probability values, and the threshold determines whether the text belongs to a specific event category, and outputs multi-category classification results;

[0040] The event extraction and coreference resolution module includes a financial event data annotation and model fine-tuning unit and a financial event extraction and coreference resolution unit; the financial event data annotation and model fine-tuning unit generates multiple types of financial event samples, annotates and saves them, and performs 10-shot supervised full-scale fine-tuning based on UIE-base;

[0041] The cross-document event interaction layer and key event extraction module include an event interaction graph construction layer and a key event screening and completion unit; the event interaction graph construction layer takes each event as a node in the graph, calculates the weight of the edge according to the similarity of the event content, and uses the Euclidean distance to calculate the similarity between events to form an edge weight reflecting the event relationship; the event interaction graph supports a hierarchical interaction mechanism, so that the event information of the same company can flow between different layers, thereby realizing the dynamic association of events;

[0042] The key event screening and completion unit screens according to the degree of event correlation, identifies and selects key events with important impact, and completes the missing information; finally, the cross-document event interaction layer and the key event extraction module output structured key event information and store it in the database.

[0043] The financial event extraction and coreference resolution unit applies the UIE framework to accurately extract financial events; during the event extraction process, the financial event extraction and coreference resolution unit first identifies the event information in the text, and aggregates similar items belonging to the same event through similarity calculation; through coreference resolution processing, the financial event extraction and coreference resolution unit block merges repeated or similar events in the announcement, completes the missing argument information, and generates complete structured event information; finally, the financial event extraction and coreference resolution unit saves the extraction results in JSON and CSV formats for subsequent analysis and processing.

[0044] The cross-document event interaction layer and key event extraction module also includes a result storage unit, which outputs the processed event information in a structured manner and stores it through a database to support subsequent analysis and query requirements.

[0045] The beneficial effects of adopting the above technical solution are:

[0046] 1. This paper proposes an event extraction method for long text financial announcements. The text summary module is used to refine the announcement content, remove redundant information, reduce processing resources and improve extraction accuracy. At the same time, a co-reference ablation strategy is proposed to effectively improve the accuracy of event recognition in the case of co-reference phenomenon in single-document event extraction of small language models.

[0047] 2. The present invention proposes a cross-document event association method based on an event interaction graph. By constructing an event interaction layer, the flow and complementarity of event information between multiple documents is realized, the integrity and coherence of event descriptions are ensured, the comprehensiveness and accuracy of key event extraction are significantly enhanced, and richer event information support is provided for financial analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of the structure of a long text financial event extraction and key event extraction system based on a small language model provided in an embodiment of the present invention;

[0049] Figure 2 This is a flow chart of the method for extracting long text financial events and key events based on a small language model adopted in the present invention. DETAILED DESCRIPTION

[0050] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0051] like Figure 1As shown, the long text financial event extraction and key event extraction system based on the small language model is as follows: the system includes a data cleaning module, a long text summary module, an event classification module, an event extraction and syntactic ablation module, a cross-document event interaction layer and a key event extraction module;

[0052] Methods include:

[0053] S1. Obtain the financial announcement data of listed companies through the data cleaning module and perform cleaning processing to generate standardized announcement titles and contents;

[0054] S2. Input the cleaned standard financial announcement text into the long text summary module to generate sentence vector representation and obtain high-importance sentences to form a complete summary;

[0055] S3, input the summary content into the event classification module for encoding, generate feature vectors, optimize model performance through the classification layer, and obtain multi-category event classification results;

[0056] S4. Input the financial announcement summary and event classification results into the event extraction and co-reference ablation module, apply the small language model after annotation and fine-tuning through the UIE framework to perform event extraction and co-reference event ablation, obtain complete structured event information and save it;

[0057] S5. Input the structured event representation into the cross-document event interaction layer and the key event extraction module to obtain the missing information completion and top-k (usually k is 5 or 10) key event information of the same listed company across documents, and store the processed structured event information in the database.

[0058] The data cleaning module is responsible for obtaining the csv format data of financial announcements of listed companies, extracting key information such as announcement title, announcement content, company name and release time, and saving it as a standard format file. At the same time, the module cleans the announcement title and content, removes non-structured symbols, tags and irrelevant characters, generates cleaned announcement title and content, and provides a standardized data basis for subsequent processing;

[0059] The long text summary module includes a text embedding and importance calculation unit, which is responsible for segmenting the cleaned announcement content into sentences and generating vectorized representations of the sentences. The module further determines the relative importance of each sentence based on a weight calculation method.

[0060] Furthermore, the long text summary module also includes a summary generation unit, which selects sentences with high document frequency, sorts them by importance, and finally generates a complete summary to highlight the core content of the text;

[0061] The event classification module is responsible for encoding the summary content and converting the text into a feature vector. This module sets the classification layer so that the model can generate classification scores for various types of events and optimize the model performance. Through normalization operations, the module converts the classification scores into probability values, thresholds to determine whether the text belongs to a specific event category, and outputs multi-category classification results;

[0062] The event extraction and syntactic ablation module includes a financial event data labeling and model fine-tuning unit, which generates multiple types of financial event samples, labels and saves them, and performs 10-shot supervised full-scale fine-tuning based on UIE-base.

[0063] Furthermore, the event extraction and syntactic ablation module includes a financial event extraction and syntactic ablation unit, which applies the UIE (Universal Information Extraction) framework to accurately extract financial events. During the event extraction process, the module first identifies the event information in the text, and aggregates similar items that may belong to the same event through similarity calculation. Through syntactic ablation processing, the module merges repeated or similar events in the announcement, completes the missing argument information, and generates complete structured event information. Finally, the module saves the extraction results in JSON and CSV formats for subsequent analysis and processing.

[0064] The cross-document event interaction layer and key event extraction module include an event interaction graph construction layer, which takes each event as a node in the graph, calculates the weight of the edge according to the similarity of the event content, and uses the Euclidean distance to calculate the similarity between events to form edge weights that reflect the relationship between events. The event interaction graph supports a hierarchical interaction mechanism that enables event information of the same company to flow between different layers, thereby realizing dynamic association of events.

[0065] The cross-document event interaction layer and key event extraction module include a key event screening and completion unit, which screens, identifies and selects key events with important impacts according to the degree of event correlation, and completes the missing information. Finally, the module outputs structured key event information and stores it in a database to provide data support for further query and analysis.

[0066] Furthermore, the cross-document event interaction layer and key event extraction module also include a result storage unit, which outputs the processed event information in a structured manner and stores it through a database to support subsequent analysis and query requirements.

[0067] On the other hand, the present invention also provides a method for extracting long text financial events and key events using the above-mentioned small language model, comprising the following steps:

[0068] Step 1: Obtain and preprocess the financial announcement dataset, including:

[0069] Step 1.1: Data acquisition and format conversion: Use a dedicated data acquisition module to crawl the financial announcement data of listed companies, extract key information (announcement title, announcement content, company name, release time, etc.), and save these contents as csv Format.

[0070] Step 1.2: Clean the announcement title and announcement content fields in csv, denoted as T title and T content , unstructured symbols, HTML tags and irrelevant characters in the text are removed, and the cleaned announcement title and announcement content are recorded as and for subsequent processing.

[0071] Step 2: Perform sentence embedding and summary generation on the data obtained in step 1, including:

[0072] Step 2.1: Sentence segmentation and embedding generation: Split into a sentence set T = {s1, s2, ..., sn} and use the bert-base-chinese model to generate sentence embeddings. For each sentence s i , use bert-base-chinese to generate a d-dimensional vector representation h i , thus obtaining the implicit representation matrix H = {h1, h2, ..., hn} of the text, where each Representative sentence s i Embedding.

[0073] Step 2.2: Sentence importance calculation: Calculate each sentence s i The importance weight w(s i ) to select representative sentences. Use TF-IDF weight w(s i ) to calculate: Among them, TF(w) represents the term frequency of word w, and IDF(w) is the inverse document frequency.

[0074] Step 2.3: Summary generation: According to the weight w(s) of each sentence i ) are sorted, and the top m important sentences are selected from high to low to form a summary candidate set And keep the original order of the sentence, and convert the candidate set S summary The sentences in are concatenated in order to form a complete summary:

[0075] Step 3: Summarize the text S summary Perform multi-label classification to classify the text into 13 financial event types C = {c1, c2, ..., c 13}, including company listing, shareholder increase, enterprise acquisition, etc. Specifically including:

[0076] Step 3.1: Feature vectorization: summary Use bert-base-chinese to encode and embed the text into a feature vector representation F = {f1, f2, ..., f d}, d is the feature dimension, and its embedded features can retain the semantic information of the event category.

[0077] Step 3.2: Multi-label classification layer setting: Introduce the multi-label classification head of the bert-base-chinese model and convert the output feature vector F = {f1, f2, ..., f d}, input to the classification head to generate the category classification score The specific calculation formula is: The weight matrix and bias term are automatically adjusted during the model training process, and the multi-label classification head is optimized using the binary cross-entropy (BCE) loss function. For each category, the predicted probability p(c i |T) and the true label y i , the BCE loss function is calculated as: When the text belongs to category c1.

[0078] Step 3.3: Normalization and threshold determination: Use the sigmoid function to map the score to probability Calculate the probability of each category independently and set a threshold τ to determine whether the text belongs to a specific category c i , the final multi-label classification result can be expressed as:

[0079] Step 4: Use the open source UIE-Base model combined with the UIE general information extraction framework proposed by ACL2022 to extract financial events in detail, and use the event co-reference resolution method to resolve and merge the co-referencing events in the same document. Specifically include:

[0080] Step 4.1: Data annotation: Provide necessary data support for event extraction. 10-shot samples were generated for 13 types of financial events, annotated on the doccano data annotation platform, and the annotation files were stored in JSON format, including event categories, trigger words, arguments, and argument positions.

[0081] Step 4.2: Fine-tune the UIE-Base model: Use the UIE module in the Python PaddleNLP library for supervised full-scale fine-tuning, use AutoTokenizer to load the tokenizer, load dataset to load the training dataset and validation dataset, and use the convert_example function to convert them to fit the model input. At the same time, set appropriate hyperparameters, such as learning rate and training rounds, for efficient fine-tuning.

[0082] Step 4.3: Financial event extraction and similarity calculation: According to the classification of events, they are divided into 13 categories of schemas. The Taskflow library under the PaddleNLP framework is applied to call the financial event extraction model based on UIE-base. In the extraction task, Taskflow is called, the parameter is set to information_extraction, and the schema is put into the model in the form of a prompt to start extraction. In the first round of extraction results, similar events (A, B, and C events) extracted from the same text are calculated for similarity to determine whether they are the same event. First, use the pre-trained model all-MiniLM-L6-v2 in the Sentence Transformers library to encode the text arguments {t1, t2, t3} into vector representations {v1, v2, v3}, v i =Encode(t1). Use cosine similarity to calculate the similarity between the encoded vectors:

[0083] Step 4.4: Merge elimination and output formatting: For similar events, first set a threshold, such as threshold = 0.8, to control the merging condition. If the similarity is higher than the threshold, it is considered as the same event chain, and a list L = {(e i , e j ,cosine_similarity(e i , e j ))}Select the event with the best expression completeness as the benchmark template base The completeness is represented by the event argument set R(e) = {r1, r2, ..., r n}, where r i is an argument. For example, R(e base ) represents the argument set of the benchmark event. If the argument r is missing in the benchmark event j , then complete it from other events and select the argument with the largest probability value as the final result, which is expressed as: ), where P(r j |e i) indicates that in event e i Zhonglun Yuan j The merged events will be stored together with the announcement type, announcement title, announcement content, company name and event time information as DataFrame = {announcement type, announcement title, announcement content, company name, event time, r1, r2, ...}.

[0084] Step 5: Design an event interaction layer for cross-document event completion and key event identification. Specifically include:

[0085] Step 5.1: Cross-document event similarity calculation and completion: Each event is regarded as a node e i , use SentenceTransformers to transform e i Coded as V i . Use the Euclidean distance to calculate the distance between any two nodes e i and e j The similarity S ij , d ij =||V i -V j ||,S ij The larger the event e i and e j All event nodes and their similarity edge weights are constructed into an event interaction graph G = (E, S), where E is the event node set and S is the edge weight matrix.

[0086] Step 5.2: Key event screening and completion: For each event node e i Calculate its relevance score R(e i )=∑ j∈N(i) S ij , where N(i) represents the node e i The higher the relevance score, the stronger the relevance of the event in the network. i ), select the first 10 event nodes to form a key event chain, denoted as set ε key For each node e in the key event chain k ∈ε key , whose arguments are the set R k = {r k,1 , r k,2 , ..., r k,m}. If the event node e k If an argument is missing from , you can find the missing argument: Where P(r j ) is the argument r jThe probability of occurrence in the adjacent nodes, N(k) represents the set of adjacent nodes with high similarity. kj The most reliable candidate completion argument is selected by the size of the adjacent node. j,m+1 To complete e k Final expression: After completing all key event nodes, output the complete event chain information.

[0087] Step 5.3: Output storage: Use Python's json library to output the processed event extraction results and key event chain information in a hierarchical JSON structure, use the pymongo library to establish a connection with MongoDB, and use the insert operation to insert the structured JSON event information into the database for efficient retrieval and management. This realizes the connection between data in JSON format and database storage to support subsequent analysis and query.

Claims

1. A method for extracting long text financial events and key events based on a small language model, characterized in that: The steps include: Step 1: Obtain and preprocess the financial announcement dataset to generate standardized announcement titles and announcement contents; Step 2: Perform sentence embedding and summary generation on the data obtained in step 1 to obtain high-importance sentences to form a complete text summary S summary ; Step 3: Summarize the text S summary Multi-label classification is performed and classified into the 13 financial event types set C = {c1, c2, ..., c 13 }, including company listing, shareholder reduction, shareholder increase, enterprise acquisition, enterprise financing, share repurchase, pledge, contact pledge, enterprise bankruptcy, loss, interview, winning bid, and senior management changes; Step 4: Fine-tune the small language model based on the UIE general information extraction framework to extract financial events in detail, and merge co-reference events in the same document through the event co-reference resolution method; Step 5: Design an event interaction layer for cross-document event completion and key event identification to obtain the missing information completion and top-k key event information of the same listed company across documents.

2. The method for extracting long text financial events and key events based on a small language model according to claim 1, characterized in that: The step 1 specifically includes: Step 1.1: Data acquisition and format conversion: crawl the financial announcement data of listed companies through the data acquisition module, extract key information, and save it as a csv file; Step 1.2: Trace the announcement title in the csv file title And the announcement content T content The fields are cleaned to remove non-structured symbols, HTML tags and irrelevant characters in the fields. The cleaned announcement title and announcement content are recorded as and 3. The method for extracting long text financial events and key events based on a small language model according to claim 2 is characterized in that: The step 2 specifically includes: Step 2.1: Sentence segmentation and embedding generation: Divide into sentence set T = {s1, s2, ..., sn}, use the pre-trained model bert-base-chinese to generate sentence embedding; for each sentence s i , use bert-base-chinese to generate a d-dimensional vector representation h i , we get the implicit representation matrix H = {h1, h2, ..., hn} of the text, where each Representative sentence s i Embedding Step 2.2: Sentence importance calculation: Calculate each sentence s i The importance weight w(s i ) to select representative sentences; use TF-IDF weight w(s i ) to calculate: Among them, TF(w) represents the word frequency of word w, and IDF(w) is the inverse document frequency; Step 2.3: Summary generation: According to the weight w(s) of each sentence i ) are sorted, and the top m important sentences are selected from high to low to form a summary candidate set And keep the original order of the sentence, and convert the candidate set S candidate The sentences in are concatenated in order to form a complete summary:

4. The method for extracting long text financial events and key events based on a small language model according to claim 3 is characterized in that: The step 3 specifically includes: Step 3.1: Feature vectorization: For the complete summary S summary Use the bert-base-chinese model for encoding and embed the text into a feature vector representation F = {f1, f2, ..., f d }, d is the feature dimension, and the feature vector represents the semantic information that can preserve the event category; Step 3.2: Multi-label classification layer setting: Introduce the multi-label classification head of the bert-base-chinese model and convert the output feature vector F = {f1, f2, ..., f d }, input to the multi-label classification head to generate the category classification score The specific calculation formula is: The weight matrix and bias term are automatically adjusted during the bert-base-chinese model training process, and the binary cross entropy loss function is used to optimize the multi-label classification head. For each category, the prediction probability p(c i |T) and the true label y i , the BCE loss function is calculated as: y i =1 when the text belongs to category c1; Step 3.3: Normalization and threshold determination: Use the sigmoid function to map the category classification scores to probabilities Calculate the probability of each category independently and set a threshold τ to determine whether the text belongs to a specific category c i , the final multi-label classification result is expressed as:

5. The method for extracting long text financial events and key events based on a small language model according to claim 4 is characterized in that: The step 4 specifically includes: Step 4.1: Data annotation: Generate 10-shot samples for 13 types of financial events, annotate them on the doccano data annotation platform, and store the annotation files in JSON format, including event categories, trigger words, arguments, and argument positions; Step 4.2: Fine-tune the UIE-Base model as a financial event extraction model: Use the UIE module in the Python PaddleNLP library for supervised full-scale fine-tuning, use AutoTokenizer to load the tokenizer from the pre-trained UIE-Base model, load_dataset to load the training dataset and validation dataset, and use the convert_example function to convert them as model input; set hyperparameters at the same time; Step 4.3: Financial event extraction and similarity calculation: According to the event classification results, they are divided into 13 categories of extraction templates, and the Taskflow library under the PaddleNLP framework is applied to call the financial event extraction model fine-tuned in step 4.2; in the extraction task, call Taskflow, set the parameter to information_extraction, and put the schema into the financial event extraction model in the form of prompt to start extraction; in the first round of extraction results, similarity calculations are performed on similar events extracted from the same text to determine whether they are the same event; first, use the pre-trained model all-MiniLM-L6-v2 in the Sentence Transformers library to encode the text arguments {t1, t2, t3} into vector representations {v1, v2, v3}, v i =Encode(t1); use cosine similarity to calculate the similarity between encoded vectors Step 4.4: Co-reference resolution and output formatting: For similar events, first set a threshold threshold to control the merging condition; if the similarity is higher than the threshold, it is considered as the same event chain, and output a list L containing similar event pairs and their similarities = {(e i , e j ,cosine_similarity(e i , e j ))}Select the event with the best expression completeness as the benchmark template base ; The completeness of the event argument set R(e) = {r1, r2, ..., r n }, where r i is an argument; the argument r is missing in the benchmark event j , then complete it from other events and select the argument with the largest probability value as the final result, which is expressed as: Where P(r j |e i ) indicates that in event e i Zhonglun Yuan j The merged events will be stored together with the announcement type, announcement title, announcement content, company name and event time as DataFrame = {announcement type, announcement title, announcement content, company name, event time, r1, r2, …}.

6. The method for extracting long text financial events and key events based on a small language model according to claim 5, characterized in that: The step 5 specifically includes: Step 5.1: Cross-document event similarity calculation and completion: Each event is regarded as a node e i , use SentenceTransformers to transform e i Coded as V i ; Use Euclidean distance to calculate the distance between any two nodes e i and e j The similarity S ij , d ij =||V i -V j ||,S ij The larger the event e i and e j The more similar they are; all event nodes and their similarity edge weights are constructed into an event interaction graph G = (E, S), where E is the event node set and S is the edge weight matrix; Step 5.2: Key event screening and completion: For each event node e i Calculate its relevance score R(e i )=∑ j∈N(i) S ij , where N(i) represents the node e i The set of adjacent nodes; the higher the correlation score, the stronger the correlation of the event in the network; according to the correlation score R(e i ), select the first k event nodes to form a key event chain, denoted as set ε key ; For each node e in the key event chain k ∈ε key , whose arguments are the set R k = {r k,1 , r k,2 ,…,r k,m }; Event node e k When an argument is missing from , find the missing argument: Among them, P(r j ) is the argument r j The probability of occurrence in the adjacent nodes, N(k) represents the set of adjacent nodes with high similarity; according to the edge weight S kj The size of the argument is used to select the most reliable candidate completion argument; the argument value r with the highest edge weight among the adjacent nodes is selected j,m+1 To complete e k After completing all key event nodes, the complete event chain information is output.

7. A long text financial event extraction and key event extraction system based on a small language model, characterized in that: The method for extracting long-text financial events and key events based on a small language model according to any one of claims 1 to 6 is adopted; the system for extracting long-text financial events and key events based on a small language model comprises a data cleaning module, a long-text summary module, an event classification module, an event extraction and coreference resolution module, a cross-document event interaction layer and a key event extraction module; S1. Obtain the financial announcement data of listed companies through the data cleaning module and perform cleaning processing to generate standardized announcement titles and announcement contents; S2. Input the cleaned and standardized announcement title and content into the long text summary module to generate sentence vector representation and obtain high-importance sentences to form a complete summary; S3, input the complete summary content into the event classification module for encoding, generate feature vectors, optimize model performance through the classification layer, and obtain multi-category event classification results; S4. Input the financial announcement summary and event classification results into the event extraction and coreference resolution module, apply the small language model after annotation and fine-tuning through the UIE framework to perform event extraction and coreference event resolution, obtain complete structured event information and save it; S5. Input the structured event information into the cross-document event interaction layer and the key event extraction module to obtain the missing information completion and top-k key event information of the same listed company across documents, and store the processed structured event information into the database.

8. The long text financial event extraction and key event extraction system based on a small language model according to claim 7 is characterized in that: The data cleaning module is responsible for obtaining the csv format data of the listed company's financial announcements, extracting key information, and saving it as a standard format file; the data cleaning module cleans the announcement title and content, removes non-structured symbols, tags and irrelevant characters, generates cleaned announcement title and content, and provides a standardized data basis for subsequent processing; The long text summary module includes a text embedding and importance calculation unit and a long text summary module; the text embedding and importance calculation unit is responsible for sentence segmentation of the cleaned announcement content and generating a vectorized representation of the sentence; The long text summary module further determines the relative importance of each sentence based on a weight calculation method; the summary generation unit selects sentences with high document frequency, sorts them by importance, and finally generates a complete summary to highlight the core content of the text; The event classification module is responsible for encoding the summary content and converting the text into a feature vector; the event classification module sets a classification layer so that the model generates classification scores for various types of events and optimizes model performance; through normalization operations, the event classification module converts the classification scores into probability values, and the threshold determines whether the text belongs to a specific event category, and outputs multi-category classification results; The event extraction and coreference resolution module includes a financial event data annotation and model fine-tuning unit and a financial event extraction and coreference resolution unit; the financial event data annotation and model fine-tuning unit first randomly selects ten samples of each type of financial event, annotates and saves them, and performs 10-shot supervised full-scale fine-tuning based on UIE-base; The cross-document event interaction layer and key event extraction module include an event interaction graph construction layer and a key event screening and completion unit; The event interaction graph construction layer takes each event as a node in the graph, calculates the edge weight according to the similarity of the event content, and uses the Euclidean distance to calculate the similarity between events to form an edge weight that reflects the event relationship; The event interaction graph supports a hierarchical interaction mechanism, which enables event information of the same company to flow between different layers, thus achieving dynamic association of events; The key event screening and completion unit screens according to the degree of event correlation, identifies and selects key events with important impact, and completes the missing information; finally, the cross-document event interaction layer and the key event extraction module output structured key event information and store it in the database.

9. The long text financial event extraction and key event extraction system based on a small language model according to claim 8 is characterized in that: The financial event extraction and coreference resolution unit applies the UIE framework to accurately extract financial events; during the event extraction process, the financial event extraction and coreference resolution unit first identifies the event information in the text, and aggregates similar items belonging to the same event through similarity calculation; through coreference resolution processing, the financial event extraction and coreference resolution unit block merges repeated or similar events in the announcement, completes the missing argument information, and generates complete structured event information; finally, the financial event extraction and coreference resolution unit saves the extraction results in JSON and CSV formats for subsequent analysis and processing.

10. The long text financial event extraction and key event extraction system based on a small language model according to claim 8, characterized in that: The cross-document event interaction layer and key event extraction module also includes a result storage unit, which outputs the processed event information in a structured manner and stores it through a database to support subsequent analysis and query requirements.

Citation Information

Patent Citations

  • Event Information Fusion Methods and Systems

    CN102298635A

  • Event element extraction method and device based on semantic analysis and prompt learning, electronic equipment and storage medium

    CN116186241A

  • Financial announcement event extraction method fusing paragraph and document features

    CN118673919A

  • Financial event and relationship extraction

    US20090327115A1

  • Generative event extraction method based on ontology guidance

    US20240143633A1

Cited By

  • Cross-language long document abstracting method based on dynamic potential key information constraint

    CN121542420A

  • Cross-language long document summarization method based on dynamic latent key information constraints

    CN121542420B