A method and system for event extraction in portfolio management combining extraction and classification tasks
Patent Information
- Application Number
- CN202310522765.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-05-10
AI Technical Summary
但目前针对资管行业的事件抽取技术还不够成熟,无法有效的在长文本组成的文档抽取事件;若事件中的论元在原文中没有出现,即无法抽取,需要进行分类时,难以处理此类任务等等
Smart Images

Figure CN116644150B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of asset management, specifically to a method and system for extracting asset management events that combines extraction and classification tasks. Background Technology
[0002] The asset management industry refers to the process by which investors entrust financial institutions to invest and manage their assets. This includes institutions such as funds, bank wealth management products, insurance companies, and trusts. Currently, the asset management industry is expanding rapidly, but it primarily relies on the individual abilities of fund managers to manage and allocate investors' assets; the integration of financial technology is still relatively insufficient. Furthermore, faced with the massive amounts of information generated in the internet age, individuals and teams struggle to quickly filter out high-value information and respond to events in a timely manner.
[0003] Event extraction technology is a key information extraction technique that extracts relevant event information from given natural language text and identifies the event's components such as time, location, and country. However, current event extraction technologies for the asset management industry are not yet mature enough. They cannot effectively extract events from documents composed of long texts; if arguments in an event do not appear in the original text, they cannot be extracted, making it difficult to handle tasks requiring classification. Furthermore, sequence labeling-based models cannot solve the problem of entity overlap, meaning that the same word may represent multiple event extraction arguments. Summary of the Invention
[0004] This application provides a method for extracting asset management events by combining extraction and classification tasks, which can be used to solve document-level event extraction tasks in the asset management field.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] A method for extracting asset management events that combines extraction and classification tasks includes the following steps:
[0007] The input document is preprocessed to obtain the processed document;
[0008] Based on the generative model, the processed document is subjected to argument extraction to obtain event argument extraction results.
[0009] The event extraction argument results are processed based on the knowledge enhancement model to obtain event classification argument results;
[0010] The processed document is then processed based on the event extraction argument results and the event classification argument results to obtain the event extraction results.
[0011] Preferably, the preprocessing method includes:
[0012] Remove extra spaces, line breaks, and non-Chinese / English symbols from the input document;
[0013] Remove sentences from the input document where Chinese characters constitute a very small percentage of the total text.
[0014] Remove excessively short and excessively long Chinese text from the input document.
[0015] Preferably, the method for extracting arguments includes:
[0016] Based on the aforementioned generative model, the processed document is converted into word embedding vectors, and prompts are constructed.
[0017] Based on the prompt, feature extraction is performed to obtain prompt features, and event argument extraction is performed based on the prompt features to obtain the event argument extraction result.
[0018] Preferably, the method for obtaining the event classification argument results includes:
[0019] Based on the event-extracted argument results, construct classification argument prompts;
[0020] The classification argument cues are processed by a knowledge enhancement model to obtain the event classification argument results.
[0021] Preferably, the method for obtaining the event extraction result includes:
[0022] The event arguments are extracted and post-processed by combining pre-trained model word embedding representation and keyword dictionary technology, and then linked with database data to obtain the first result;
[0023] Extract all events from the first result and determine the event type of each event;
[0024] Events of the same type are merged and completed to obtain the event extraction result.
[0025] This application also provides an asset management event extraction system that combines extraction and classification tasks, including: a document preprocessing module, an argument extraction module, an argument classification module, and an entity linking module;
[0026] The document preprocessing module is used to preprocess the input document to obtain the processed document;
[0027] The argument extraction module is used to extract arguments from the processed document based on the generative model to obtain event argument extraction results.
[0028] The classification argument module is used to process the event extraction argument results based on the knowledge enhancement model to obtain event classification argument results;
[0029] The entity connection module is used to process the processed document based on the event extraction argument results and the event classification argument results to obtain the event extraction results.
[0030] Preferably, the workflow of the document preprocessing module includes:
[0031] Remove extra spaces, line breaks, and non-Chinese / English symbols from the input document;
[0032] Remove sentences from the input document where Chinese characters constitute a very small percentage of the total text.
[0033] Remove excessively short and excessively long Chinese text from the input document.
[0034] Preferably, the workflow of the argument extraction module includes:
[0035] Based on the aforementioned generative model, the processed document is converted into word embedding vectors, and prompts are constructed.
[0036] Based on the prompt, feature extraction is performed to obtain prompt features, and event argument extraction is performed based on the prompt features to obtain the event argument extraction result.
[0037] Preferably, the workflow of the classification argument module includes:
[0038] Based on the event-extracted argument results, construct classification argument prompts;
[0039] The classification argument cues are processed by a knowledge enhancement model to obtain the event classification argument results.
[0040] Preferably, the workflow of the entity linking module includes:
[0041] The event arguments are extracted and post-processed by combining pre-trained model word embedding representation and keyword dictionary technology, and then linked with database data to obtain the first result;
[0042] Extract all events from the first result and determine the event type of each event;
[0043] Events of the same type are merged and completed to obtain the event extraction result.
[0044] The beneficial effects of this application are as follows:
[0045] This application can simultaneously handle extracted and classified arguments in events, bringing greater flexibility to the definition of event templates to suit a wider range of application scenarios. Furthermore, due to the application of a generative model, it can effectively solve the problem of entity overlap. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of this application, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a schematic diagram of the method flow of Embodiment 1 of this application;
[0048] Figure 2 This is a schematic diagram of a specific event template for Embodiment 1 of this application;
[0049] Figure 3 This is a schematic diagram of the system structure of Embodiment 2 of this application. Detailed Implementation
[0050] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0051] First, let's introduce the technical terms used in this embodiment:
[0052] An extracted argument is defined as an argument that is relevant to an event and whose start and end positions can be determined in the original text. For example, in the event "price increase" in "natural gas prices are expected to double in the second half of this year", the "product" argument is an extracted argument, and its content is "natural gas". The argument starts at position 5 and ends at position 8.
[0053] A classification argument is defined as an argument that is related to an event but whose start and end positions cannot be determined in the original text, and which needs to be further classified using classification methods. For example, in the event of "price increase" in "natural gas prices are expected to double in the second half of this year", the argument "REALIS" (which determines whether an event has already occurred or is expected to occur in the future) is a classification argument, and its classification is expected to occur in the future.
[0054] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] Example 1
[0056] In this first embodiment, as Figure 1 As shown, an asset management event extraction method combining extraction and classification tasks includes the following steps:
[0057] S1. Preprocess the input document to obtain the processed document.
[0058] The preprocessing methods include: removing redundant spaces, line breaks, and non-Chinese / English symbols from the input document; removing sentences in the input document where Chinese characters constitute a small proportion of the overall text; and removing excessively short or long Chinese text from the input document.
[0059] In this embodiment, 13 event templates in the asset management field are first defined. Each event template includes an event name and the event elements to be extracted. Specific event templates are as follows: Figure 2 As shown; then the input document Text is preprocessed to remove extra spaces, line breaks, and non-Chinese / English symbols; sentences in which Chinese characters account for less than 30% of the total text are removed; and excessively short and excessively long Chinese texts are removed, where excessively short texts are those with fewer than 3 characters and excessively long texts are those with more than 512 characters.
[0060] S2. Based on the generative model, extract arguments from the processed document to obtain the event-extracted argument results.
[0061] The method for extracting arguments includes: based on the generative model, converting the processed document into word embedding vectors and constructing prompts; extracting features based on the prompts to obtain prompt features, and extracting arguments based on the prompt features to obtain the event extraction argument results.
[0062] In this embodiment, a prompt input generation model is constructed for the prompt generation model. Specifically, the construction method is as follows: assuming x is an event name, and s... i ∈{s0,s1,……,s n Let} represent the extracted event arguments for event x, and n represent the total number of extracted arguments for that event. Then, prompt_ext is constructed as follows:
[0063] [event]x[ / event][asso_type]s0[ / asso_type]……[asso_type]s n [ / asso_type]
[0064] Here, `event` represents the event, and `asso_type` represents the entity type. After constructing the prompt, the original text is appended to it, and any text exceeding the length is truncated. The final input `x_ext` for the model is as follows:
[0065] [Prompt_ext]+[Text]
[0066] The text is then input into a T5 generation model fine-tuned for asset management events for prediction. The T5 generation model consists of an encoder and a decoder. The encoder is responsible for encoding the semantic features of the text, and the decoder is used to generate the extraction results. The specific structure of the model is as follows:
[0067] The encoder consists of 12 sequentially stacked coding blocks. Each coding block contains a multi-head attention network layer and a feedforward neural network layer. Each layer applies residuals to avoid the vanishing and exploding gradient problems of deep neural networks. Each multi-head attention network consists of 16 self-attention networks.
[0068] The decoder consists of 12 sequentially stacked decoding blocks. Each decoding block contains a masked multi-head attention layer, a multi-head attention layer, and a feedforward neural network layer, with each layer also employing residual techniques. The masked multi-head attention layer adds a masking mechanism to the multi-head attention layer to prevent information leakage during training. The key and value of the multi-head attention layer in the decoding block are the encoder's feature vectors.
[0069] After inputting text into the generative model, the autoregressive model generates a structured output, y1, as shown below:
[0070] y1 = [X trigger words: [s0: ..., s1: ..., ..., s n :……]).
[0071] S3. Process the event extraction argument results based on the knowledge enhancement model to obtain the event classification argument results.
[0072] The method for obtaining event classification argument results includes: constructing classification argument prompts based on event extraction argument results; and classifying the classification argument prompts using a knowledge enhancement model to obtain event classification argument results.
[0073] In this embodiment, the prompt_cls for processing classification arguments will be constructed by combining the output result y1 and the original text Text. The specific construction method is as follows:
[0074] Construct prompt_cls by filling the event name X, event trigger word Trigger, and classification argument Cls from y1 into the following template in the form of slots:
[0075] Event: [X], Event Trigger: [Trigger], Event Element: [Cls]
[0076] Then, prompt_cls is concatenated with the original text Text and used as the input x_cls to the knowledge model:
[0077] [prompt_cls]+[Text]
[0078] Then, for each argument ci that needs to be classified, an x_cls is constructed as input and the Ernie knowledge augmentation model, fine-tuned for asset management events, is used for classification. Finally, each c... i The classification results are added to y1 above to form y2:
[0079] y2 = [X trigger words: [s0: ..., s1: ..., ..., s n :...,c1:...,...,c m :……]]
[0080] Among them, c i Let represent the i-th classification argument, and m represent the total number of classification arguments for event X.
[0081] S4. Based on the event extraction argument results and the event classification argument results, process the processed document to obtain the event extraction results.
[0082] The method for obtaining the event extraction results includes: post-processing the event extraction arguments by combining pre-trained model word embedding representation and keyword dictionary technology, and linking them with database data to obtain the first result; extracting all events from the first result and determining the event type of each event; merging and completing events of the same event type to obtain the event extraction results.
[0083] In this embodiment, all extracted arguments {s0, ..., s} in y2 are... n Matching with the standard asset management database, each argument s i The specific matching method is as follows:
[0084] First, s i Perform a full-word match against a pre-constructed keyword dictionary. If a match is successful, then... i Change to the dictionary word s after matching i If whole-word matching fails, the pre-trained language model text2vec is used to obtain the s. i The word embedding vectors are compared with the similarity of all words in the mapping list of the database, and the word with the highest similarity is taken as s. i The specific similarity calculation formula is as follows:
[0085]
[0086] Where x i s i Each feature value of the word embedding vector, y i This represents the feature values of the word embedding vectors mapped to the database. Then, s... i 'Replace s i , forming y3:
[0087] y3 = [X trigger words: [s0': ..., s1': ..., ..., s n ':...,c1:...,...,c m :……]]
[0088] Extract all events from y3. For events of the same type, merge them according to the set merging method. The specific merging method is as follows: Extract an event list elist = {e1, e2, ..., e6} based on the event name X, where ei represents the i-th instance of event X extracted from the text. For each ei, compare it sequentially with other ej that belong to elist but ei ≠ ej. The comparison method is to compare event arguments one by one. If the overlap between ei and ej is higher than a preset threshold, then ei and ej are determined to be the same event. Add the unique arguments in ej to ei, keep ei, and delete ej.
[0089] This system identifies key entities in news headlines using an ERNIE-based named entity recognition model. Key entities defined in this system represent crucial subjects related to the asset management field in the headline, such as company names, country names, industry names, and product names. Then, the system iterates through the extracted events, filling in missing event arguments for each event according to a pre-defined event template, ultimately obtaining document-level event extraction results.
[0090] Example 2
[0091] In this second embodiment, as Figure 3 As shown, an asset management event extraction system that combines extraction and classification tasks includes: a document preprocessing module, an argument extraction module, an argument classification module, and an entity linking module.
[0092] The document preprocessing module is used to preprocess the input document to obtain the processed document. The workflow of the document preprocessing module includes: removing redundant spaces, line breaks, and non-Chinese and non-English symbols from the input document; removing sentences in the input document where Chinese characters account for too small a proportion of the overall text; and removing excessively short and excessively long Chinese text from the input document.
[0093] The argument extraction module is used to extract arguments from the processed document based on the generative model to obtain event argument extraction results. The workflow of the argument extraction module includes: converting the processed document into word embedding vectors based on the generative model and constructing prompts; extracting features from the prompts to obtain prompt features, and extracting event arguments based on the prompt features to obtain event argument extraction results.
[0094] The classification argument module is used to process the event extraction argument results based on the knowledge enhancement model to obtain the event classification argument results. The workflow of the classification argument module includes: constructing classification argument prompts based on the event extraction argument results; classifying the classification argument prompts through the knowledge enhancement model to obtain the event classification argument results.
[0095] The entity linking module is used to process the processed document based on the event extraction argument results and the event classification argument results to obtain the event extraction results. The workflow of the entity linking module includes: post-processing the event extraction arguments by combining pre-trained model word embedding representation and keyword dictionary technology, and linking them with database data to obtain the first result; extracting all events in the first result and determining the event type of each event; merging and completing events of the same event type to obtain the event extraction results.
[0096] The embodiments described above are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made to the technical solutions of this application by those skilled in the art without departing from the spirit of this application shall fall within the protection scope defined by the claims of this application.
Claims
1. A method for extracting asset management events by combining extraction and classification tasks, characterized in that, Includes the following steps: The input document is preprocessed to obtain the processed document; Based on the generative model, the processed document is subjected to argument extraction to obtain event argument extraction results. The event extraction argument results are processed based on the knowledge enhancement model to obtain event classification argument results; The processed document is processed based on the event extraction argument results and the event classification argument results to obtain the event extraction results; The method for extracting arguments includes: Based on the aforementioned generative model, the processed document is converted into word embedding vectors, and prompts are constructed. Based on the prompt, feature extraction is performed to obtain prompt features, and event arguments are extracted based on the prompt features to obtain the event argument extraction results. The methods for obtaining the event classification argument results include: Based on the event-extracted argument results, construct classification argument prompts; The classification argument cues are classified using a knowledge enhancement model to obtain event classification argument results. The T5 generative model consists of an encoder and a decoder. The encoder is responsible for encoding the semantic features of the text, and the decoder is used to generate the extraction results. The model includes: The encoder consists of 12 sequentially stacked coding blocks. Each coding block contains a multi-head attention network layer and a feedforward neural network layer. Each layer applies residuals to avoid gradient vanishing and exploding in deep neural networks. Each multi-head attention network consists of 16 self-attention networks. The decoder consists of 12 decoder blocks stacked sequentially. Each decoder block contains a masked multi-head attention layer, a multi-head attention layer, and a feedforward neural network layer. Each layer also uses residual technology. The masked multi-head attention layer adds a masking mechanism on top of the multi-head attention layer. The key and value of the multi-head attention layer of the decoder block are the feature vectors of the encoder. For each argument ci that needs to be classified, an x_cls is constructed as input and the Ernie knowledge enhancement model, which has been fine-tuned by asset management events, is used for classification.
2. The asset management event extraction method combining extraction and classification tasks according to claim 1, characterized in that, The preprocessing method includes: Remove extra spaces, line breaks, and non-Chinese / English symbols from the input document; Remove sentences from the input document where Chinese characters constitute a very small percentage of the total text. Remove excessively short and excessively long Chinese text from the input document.
3. The asset management event extraction method combining extraction and classification tasks according to claim 1, characterized in that, The methods for obtaining the event extraction results include: The event arguments are extracted and post-processed by combining pre-trained model word embedding representation and keyword dictionary technology, and then linked with database data to obtain the first result; Extract all events from the first result and determine the event type of each event; Events of the same type are merged and completed to obtain the event extraction result.
4. An asset management event extraction system that combines extraction and classification tasks, characterized in that, include: Document preprocessing module, argument extraction module, argument classification module, and entity linking module; The document preprocessing module is used to preprocess the input document to obtain the processed document; The argument extraction module is used to extract arguments from the processed document based on the generative model to obtain event argument extraction results. The classification argument module is used to process the event extraction argument results based on the knowledge enhancement model to obtain event classification argument results; The entity linking module is used to process the processed document based on the event extraction argument results and the event classification argument results to obtain the event extraction results; The workflow of the argument extraction module includes: Based on the aforementioned generative model, the processed document is converted into word embedding vectors, and prompts are constructed. Based on the prompt, feature extraction is performed to obtain prompt features, and event arguments are extracted based on the prompt features to obtain the event argument extraction results. The workflow of the classification argument module includes: Based on the event-extracted argument results, construct classification argument prompts; The classification argument cues are classified using a knowledge enhancement model to obtain event classification argument results. The T5 generative model consists of an encoder and a decoder. The encoder is responsible for encoding the semantic features of the text, and the decoder is used to generate the extraction results. The model includes: The encoder consists of 12 sequentially stacked coding blocks. Each coding block contains a multi-head attention network layer and a feedforward neural network layer. Each layer applies residuals to avoid gradient vanishing and exploding in deep neural networks. Each multi-head attention network consists of 16 self-attention networks. The decoder consists of 12 decoder blocks stacked sequentially. Each decoder block contains a masked multi-head attention layer, a multi-head attention layer, and a feedforward neural network layer. Each layer also uses residual technology. The masked multi-head attention layer adds a masking mechanism on top of the multi-head attention layer. The key and value of the multi-head attention layer of the decoder block are the feature vectors of the encoder. For each argument ci that needs to be classified, an x_cls is constructed as input and the Ernie knowledge enhancement model, which has been fine-tuned by asset management events, is used for classification.
5. The asset management event extraction system combining extraction and classification tasks according to claim 4, characterized in that, The workflow of the document preprocessing module includes: Remove extra spaces, line breaks, and non-Chinese / English symbols from the input document; Remove sentences from the input document where Chinese characters constitute a very small percentage of the total text. Remove excessively short and excessively long Chinese text from the input document.
6. The asset management event extraction system combining extraction and classification tasks according to claim 4, characterized in that, The workflow of the entity linking module includes: The event arguments are extracted and post-processed by combining pre-trained model word embedding representation and keyword dictionary technology, and then linked with database data to obtain the first result; Extract all events from the first result and determine the event type of each event; Events of the same type are merged and completed to obtain the event extraction result.
Citation Information
Patent Citations
Document processing method and device, electronic equipment and storage medium
CN115130435A
Open information extraction
US20140032209A1