An Event Extraction Method, System and Storage Medium for Notes to Financial Statements

By applying the event extraction method based on the Transformer encoder in the notes to the financial statements, combining chapter-level semantic information and event theories information, the problem of inaccurate extraction of financial related events in the existing technology is solved, and a more efficient quality analysis of financial status is achieved.

CN116150361BActive Publication Date: 2025-07-01JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211680822.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-07-01
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively extract financial-related events in notes to financial statements, especially when dealing with the complex situations in the notes to financial statements, the distribution of event records is discontinuous, and the distribution of event arguments in multiple sentences, resulting in inaccurate classification of event categories, incomplete extraction of event records, and inaccurate identification of event arguments.

Method used

The Transformer encoder-based method is adopted, combining chapter-level semantic information and event theories information to identify event categories in the notes of financial statements, and simultaneous extraction of multiple event records is achieved through event table filling. Specific steps include: data preprocessing, identification and annotation of titles and paragraphs, segmentation and clauses, identification and annotation of event arguments, classification of event categories and filling of event tables.

Benefits of technology

It improves the accuracy of the event extraction of notes to financial statements, can more effectively capture chapter-level context semantics, and improves the efficiency of financial status quality analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150361B_ABST
    Figure CN116150361B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and storage medium for event extraction of financial statement footnotes. The method includes the following steps: obtaining a financial report PDF document, and obtaining a TXT document of the financial statement footnote text after data preprocessing; identifying and annotating the titles, their levels and paragraphs of the TXT document of the financial statement footnote text to obtain a title set and a paragraph set; identifying and annotating the event arguments of the financial events in the financial statement footnote based on a Transformer encoder, and simultaneously obtaining the vector representations of the event arguments; representing the semantic features of the paragraphs, titles and their levels in vectors, and splicing the vector representations of the words included in the event arguments and the vector representations of the titles and their levels into a vector matrix; learning the features of the event arguments, titles and their levels to judge the event categories, learning the features of the event arguments, titles and their levels and the memory vectors, and filling the event arguments into the current event roles of the event table based on a Transformer encoder and a linear binary classifier to obtain all the event records included in the current paragraph. The present invention extracts the titles and their levels in the financial statement footnote text as the key discourse-level semantic information of the financial statement footnote, uses the discourse-level semantic information and event argument information to identify the event categories in the financial statement footnote text, and designs a method for filling the event table to achieve the simultaneous extraction of multiple event records, thereby improving the accuracy of event extraction of financial statement footnotes as a whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of event extraction of financial statement footnotes, and particularly relates to a method, a system and a storage medium for event extraction of financial statement footnotes based on discourse structure recognition. Background Art

[0002] The quality analysis of an enterprise's financial situation is based on financial statements. By leveraging the implicit correlation between the items listed in the financial statements and the financial statement footnotes, further analysis is conducted on the financial-related events disclosed in the financial statement footnotes to obtain the quality analysis results of the enterprise's financial situation.

[0003] Due to the increasingly complex content disclosed in financial statement footnotes, when decision-makers conduct quality analysis of the financial situation, obtaining the financial events in the footnotes corresponding to the items listed in the financial statements requires a large amount of time and labor costs. Therefore, there is an urgent need for an enterprise to have a method to enable the automatic structured representation of unstructured text in financial statement footnotes, thereby improving the analysis efficiency.

[0004] The task of event extraction is to extract event arguments from unstructured text and organize them into a structured form (such as an event table), including two subtasks: identifying event categories and filling event arguments. An event table is a storage method, a two-dimensional table composed of event categories and event roles, used to describe events. The rows represent event categories, and the columns are event roles. An event argument refers to the participants or attributes of an event. An event role refers to the specific semantic role played by an event argument in an event. An event record is composed of event arguments and covers a sequence of event roles necessary for a certain event category. Each row in the event table is an event record.

[0005] Among the existing event extraction methods, the pattern matching-based methods are mainly based on syntax trees or regular expressions, which rely on event templates in specific fields and require strong professional knowledge to manually construct; the machine learning-based methods usually convert the event extraction problem into a classification problem. Common classification algorithms include logistic regression, naive Bayes, nearest neighbor, decision tree, and support vector machine, etc. They rely on large-scale annotated corpora and appropriate feature selection, have good generalization, but have great limitations in mining complex nonlinear relationships; the deep learning-based methods, typically based on convolutional networks, attention mechanisms, pre-trained models, and graph neural networks, have strong nonlinear expression capabilities and solve problems such as artificially designed features, poor scalability, and reliance on complex NLP tools. However, the classic deep learning methods focus on the semantic learning of a single sentence and can only capture short-distance sentence-level contextual semantics. Existing research has not retrieved event extraction methods for the features of financial statement notes. The event features of financial statement notes are obvious, such as there are often multiple categories of events in a note text, multiple event records of an event category are distributed in different parts of the text, and the event arguments of an event may be distributed in multiple discontinuous sentences. Existing methods easily miss long-range text-level contextual semantics, resulting in inaccurate event category classification, incomplete event record extraction, and inaccurate event argument identification.

[0006] Therefore, how to efficiently extract financial-related events based on the characteristics of financial statement notes is a technical problem that needs to be solved urgently in the field of financial status quality analysis. Summary of the invention

[0007] In order to overcome the defects and shortcomings of the prior art, the present invention provides an event extraction method for financial statement notes. First, the title and its hierarchy in the financial statement notes text are extracted as key chapter-level semantic information. Then, based on the Transformer encoder, event arguments of financial events in the financial statement notes are identified and labeled. The chapter-level semantic information and event argument information are used to identify the event category in the financial statement notes text. Finally, an event table filling method is designed to realize the simultaneous extraction of multiple event records, thereby improving the accuracy of financial statement notes event extraction as a whole.

[0008] A second object of the present invention is to provide an event extraction system for financial statement notes.

[0009] A third object of the present invention is to provide a computer-readable storage medium.

[0010] In order to achieve the above object, the present invention adopts the following technical solutions:

[0011] The present invention provides a method for extracting events from financial statement notes, comprising the following steps:

[0012] Obtain the PDF document of the financial report in the database file, convert the PDF document into a TXT document after data preprocessing, and match the TXT document of the financial statement notes text with the regular expression in the knowledge base;

[0013] Based on the knowledge base, the titles, levels and paragraphs of the TXT documents of the financial statement notes are identified and annotated to obtain the title set and paragraph set;

[0014] The paragraph set is divided into sentences and words to obtain a word list. The semantics in the paragraph is learned based on the Transformer encoder. The vector matrix of the output layer of the Transformer encoder is input into the CRF model to identify and label the event arguments of financial events in the notes to the financial statements and obtain the vector representation of the event arguments.

[0015] The vector representations of the words contained in the event argument and the vector representations of the title and its hierarchy are concatenated into a vector matrix, and the concatenated vector matrix is ​​input into the Transformer encoder to obtain a vector matrix that integrates the title, title hierarchy, and paragraph semantics, and the vector matrix is ​​input into the event category classifier to obtain the probabilities of all event categories in the current vector matrix, and the event category with the maximum probability is selected as the event category of the current vector matrix. The event role of the currently triggered event category in the predefined event information is queried through the index, and the event roles are output in a predefined order to obtain the vector representation of the event roles;

[0016] Construct a memory vector for recording the event argument filling process, concatenate the vector representation of the event argument, the vector representation of the event role and the memory vector, and input them into the Transformer encoder. Connect the output layer of the Transformer encoder to the linear binary classifier to obtain the probability of the event argument filling the current event role, select the event argument with the set probability to fill it into the current event role in the event table, repeat the iteration until all event roles of the current event category are filled, update the memory vector of the filled event argument and the vectorized representation of the event role, and obtain all event records contained in the current paragraph.

[0017] As a preferred technical solution, data preprocessing includes deleting headers, page numbers, tables and special symbols.

[0018] As a preferred technical solution, the TXT document of the financial statement notes text is matched with the regular expression in the knowledge base file, and the regular expression is expressed as:

[0019] start_line = [**Notes to financial statements|[Company|Our company|Enterprise|Group|Our Group]Basic information] and end_line = [**Documents for reference**];

[0020] Among them, [] represents a set of regular expressions, ** represents other strings, | represents an optional item, start_line represents the regular expression of the start line of the financial statement footnote text, and end_line represents the regular expression of the end line of the financial statement footnote text;

[0021] When the target string contains the rule represented by the regular expression, the re.search() function returns an object. When the target string does not contain the rule represented by the regular expression, re.search() returns None;

[0022] Traverse the preprocessed TXT document of the data. Let the string format of the current line be line. When the return value of re.search(start_line, line) is not None, that is, the expression in start_line is matched in line, start and continue to retain the current line until the return value of re.search(end_line, line) is not None to stop traversing, and obtain the TXT document of the financial statement footnote text.

[0023] As a preferred technical solution, identify and label the title, as well as the title hierarchy and paragraphs of the TXT document of the financial statement footnote text, specifically including:

[0024] Obtain the TXT document of the financial statement footnote text, obtain the regular expression for identifying the title and the set of label styles of the title in the knowledge base, and add a marking symbol at the beginning of the line identified as the title;

[0025] Traverse the TXT document of the financial statement footnote text, and judge whether the line with the marking symbol added is a sentence with complete semantics based on the binary statistical language model. If it is a sentence with complete semantics, retain the marking symbol, otherwise delete the marking symbol at the beginning of the line to obtain the candidate title;

[0026] Traverse the TXT document of the financial statement footnote text, sort out the label styles and the label order of the lines with the marking symbol added, load the set of label styles of the title from the knowledge base, each label style corresponds to a unique code, construct a hierarchical stack of labels to record the title hierarchy corresponding to each title, and label the title with a digital code containing the title hierarchy information;

[0027] Sort the text between titles in the TXT document into one line by paragraph, and label it with a digital code before the paragraph. Obtain the meaning of the code, and the title hierarchy corresponding to the paragraph is obtained from the label of the title. Based on the title and its hierarchy, combined with Chinese punctuation marks and serial number symbols, divide the text into paragraphs again.

[0028] As a preferred technical solution, obtain the regular expression for identifying titles in the knowledge base and the set of label styles for titles. The specific rules are as follows:

[0029] The first rule: It contains Chinese characters, and all Chinese character matches use Unicode encoding;

[0030] The second rule: If it contains label styles, the set of label styles is [(({0})|({1})|({0})|({1})|{0}、|{1}、|{0}.|{1}.|{0}.|{1}.|({0}).|({1}).|({0}).|({1}).|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0})、|({1})、|({0})、|({1})、|Section {0}|Section {1}|Chapter {0}|Chapter {1})], where [] represents a set, | represents a union, {0} represents an array character of all Chinese numerals, and {1} represents a numeric character of all positive integers;

[0031] The third rule: If there are no sentence features, the sentence does not include the symbols [《。!?;'];

[0032] The fourth rule: There are no data features, the sentence does not contain long data, and it can contain years;

[0033] The fifth rule: It does not end with a conjunction;

[0034] Lines that satisfy the first rule to the fifth rule are initially identified as titles. Represent the first rule and the second rule as regular expressions and store them in a list, denoted as L1. Represent the third rule, the fourth rule, and the fifth rule as regular expressions and store them in a list, denoted as L2. Traverse the TXT document of the financial statement notes. Let the string format of the current line be line. When the return value of re.search(L1, line) is not None and the return value of re.search(L2, line) is None, the current line is marked as a title. re.search is the search function in the re library of the Python tool.

[0035] As a preferred technical solution, construct a hierarchical stack of labels to record the title level corresponding to each title, and label the title with a numeric code containing title level information, specifically including:

[0036] Construct a dictionary TP. The dictionary TP contains multiple key-value pairs. A key-value pair consists of a key and a value, connected by a colon. Among them, the key is the code of the label style, and the value is the label style. There is a unique correspondence between the key and the value. Construct a dictionary T to store the title and the marked level;

[0037] Obtain the label style and serial number of the current line, compare the obtained label style with all the label styles in the dictionary TP. If there is a similar style, return the key corresponding to the value, that is, the code. If there is no similar style, judge whether the obtained serial number n is 1. If the serial number is 1, record the current label style as the value in the dictionary TP and assign a new key. If the serial number is not 1, delete the current title;

[0038] Input the code and serial number of the label style into the hierarchical stack. Judge whether the stack is empty. If it is empty, judge whether the element to be pushed onto the stack is a title with a serial number of 1. If it is correct, push it onto the stack. If it is not empty, obtain the top element of the stack and compare the element to be pushed onto the stack with the top element of the stack in terms of code. Judge the sequence relationship of the titles according to the code and serial number, and push the element to be pushed onto the stack according to the relationship between the titles;

[0039] Output the codes and serial numbers of the titles in the hierarchical stack in sequence as the values of the keys, and use the content of the titles as the values, and store them in the dictionary T to complete the extraction of the titles and their levels;

[0040] Represent the titles and their levels in the dictionary T with digital codes, and mark the current title and save it to the TXT document.

[0041] As a preferred technical solution, perform sentence splitting and word segmentation on the paragraph set to obtain a word segmentation list, specifically including:

[0042] Perform word segmentation on the current line T i to obtain the list [w1,…,w i ,…,w n , where w i represents the string of the i-th word. Input the list T i =[w1,…,w i ,…,w n into the Doc2evc model to obtain the vector representation T i —V;

[0043] Perform sentence splitting and word segmentation on the current paragraph, add the SEP label between sentences, and obtain the word segmentation list P i =[w1,…,w i ,…,SEP…,w n ;

[0044] Learn the semantics in the paragraph based on the Transformer encoder. Specifically, use the self-attention mechanism of the Transformer encoder to capture the context information of each word in the sentence. After passing through the self-attention mechanism module, the input word segmentation list obtains a weighted feature vector to obtain a vector matrix;

[0045] Based on the CRF model, obtain the conditional probability value of each sample output as the corresponding label, and output the list of annotation labels corresponding to the vector matrix.

[0046] As a preferred technical solution, input the spliced vector matrix into the Transformer encoder to obtain a vector matrix that fuses the title, title hierarchy, and paragraph semantics, denoted as: E = [e1, e2…e i ,…,e n , where e i represents the vector representation of the i-th event argument after integrating the title and its hierarchical information;

[0047] Connect the output layer of the Transformer encoder to the event classifier for event category classification, obtain the event category triggered by the current paragraph, take the maximum probability to obtain the label of the event category triggered by the current paragraph, and input the vector matrix E = [e1, e2…e i ,…,e n into the softmax classifier to calculate where W and b represent learnable parameter matrices, and k represents the number of event category labels;

[0048] Use cross-entropy to represent the information difference between the label of the true event type and the prediction result output by the softmax classifier, define the loss function as the average cross-entropy, randomly initialize W and b, and update using gradient descent.

[0049] An event extraction system for financial statement footnotes, including: a document acquisition unit, a discourse structure recognition unit, an event argument recognition unit, an event category classification unit, and an event table filling unit;

[0050] The document acquisition unit is used to obtain the PDF document of the financial report in the database file, convert the PDF document into a TXT document after data preprocessing, and combine the regular expressions in the knowledge base to match the TXT document of the financial statement footnote text;

[0051] The discourse structure recognition unit is used to identify and annotate the titles of the TXT document of the financial statement footnote text, as well as the title hierarchy and paragraphs based on the knowledge base, obtain the title set and paragraph set, represent the titles and their hierarchies in the title set as vectors based on the Doc2evc model, and input them into the event category classification unit;

[0052] The event argument recognition unit is used to split the sentences and words in the paragraph set to obtain a word segmentation list, learn the semantics in the paragraph based on the Transformer encoder, input the vector matrix of the output layer of the Transformer encoder into the CRF model, identify and annotate the event arguments of the financial events in the financial statement footnote, and obtain the vector representation of the event arguments;

[0053] The event category classification unit is used to splice the vector representations of the words included in the event argument, the title and its hierarchy into a vector matrix, input the spliced vector matrix into the Transformer encoder to obtain a vector matrix that fuses the title, the title hierarchy and the paragraph semantics, input the vector matrix into the event category classifier to obtain the probabilities of all event categories of the current vector matrix, select the event category with the highest probability as the event category of the current vector matrix, query the event roles of the currently triggered event category in the predefined event information through indexing, and output the event roles in the predefined order to obtain the vector representation of the event roles;

[0054] The event table filling unit is used to splice the vector representation of the event argument, the vector representation of the event role and the memory vector used to record the event argument filling process and input them into the Transformer encoder, connect the output layer of the Transformer encoder to a linear binary classifier to obtain the probability of the event argument filling the current event role, select the event argument with the probability being the set value and fill it into the current event role of the event table, repeat the iteration until all event roles of the current event category are filled, update the memory vector of the filled event argument and the vectorized representation of the event role, and obtain all event records included in the current paragraph.

[0055] A computer-readable storage medium stores a program, and when the program is executed by a processor, it implements the event extraction method of the financial statement note as described above.

[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0057] (1) The present invention uses regular expressions and a binary language model to screen the titles in the financial statement note text, uses a hierarchical stack to organize the hierarchical structure of the titles, and labels the titles in a digital coding manner, constructs an identification and annotation mechanism for the titles and their hierarchies in the financial statement note text, solves the problem that the event extraction method misses the long-distance discourse-level context semantics when learning the features of the financial statement note, and improves the accuracy of the event extraction method of the financial statement note when completing the event category classification task.

[0058] (2) The present invention uses the Doc2evc model to convert the context semantic feature information of the title and its hierarchy into vector representations, and uses the self-attention mechanism to convert the context semantic feature information of the words and sentences in the paragraph into vector representations, enhances the ability of the event extraction method to mine the context semantic information of the paragraphs in the financial statement note text, and uses the technical solution of vector splicing to simplify the fusion method of the paragraph semantic information and the discourse-level semantic information in the financial statement note text, and improves the feature fusion efficiency of the financial statement note text.

[0059] (3) The present invention learns event arguments, titles and their hierarchies, as well as the features of memory vectors in the financial statement note text through a Transformer encoder. The memory vector records the classification results of event arguments. On this basis, a linear binary classifier is added to determine whether the event argument matches the event role. The Transformer encoder and the linear binary classifier are used to transform the event table filling problem into a binary classification problem of whether the event argument matches the target event role, solving the problem that the same event category has multiple event records and further improving the accuracy of extracting event arguments by the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic flow chart of the event extraction method for the financial statement notes of the present invention;

[0061] Figure 2 It is a schematic diagram of the PDF document of the financial report of the present invention;

[0062] Figure 3 It is a schematic flow chart of the process of labeling tags of the present invention;

[0063] Figure 4 It is a schematic flow chart of the process of sorting out the label styles of the present invention;

[0064] Figure 5 It is a schematic flow chart of the event extraction system for the financial statement notes of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] Embodiment 1

[0067] This embodiment provides an event extraction method for financial statement notes, including the following steps:

[0068] S1: Obtain the PDF document of the financial report in the database file. After data preprocessing, convert the PDF document into a TXT document and delete the text that is not the financial statement note text;

[0069] S11: Use the pdfplumber library in the python tool to obtain the PDF document of the financial report from the database file, as Figure 2 shown;

[0070] S12: Use the extract_text function in the page module of the pdfplumber library to convert the PDF document into a TXT document, delete the header, page numbers, tables, and special symbols, and combine with the regular expressions in the knowledge base file to match the TXT document of the financial statement notes text.

[0071] (1) Delete the header: The extract_text function in the page module of the pdfplumber library reads the PDF document page by page. The first line of each page is the header and is not saved. Starting from the second line, it is saved line by line as a TXT document.

[0072] (2) Delete the page numbers: The extract_text function in the page module of the pdfplumber library reads the PDF document page by page. The last line of each page is the page number and is not saved..

[0073] (3) Delete the tables: Create a list LT to store the text strings in the tables. Use the extract_tables function in the page module of the pdfplumber library to obtain the text extracted from all the tables in the PDF document, convert the values in the tables into string format, and store them line by line in the list LT. Traverse each line of the TXT document and use the comparison operator "==" in the Python tool to determine whether the current line is equal to the values in the list LT. If the string of the current line is equal to the value in the list LT, delete the current line; if not, keep the current line.

[0074] (4) Delete special symbols: Create a list LS to store the Unicode codes of Chinese characters, numbers, English, and Chinese and English punctuation marks. Traverse each line of the TXT document and use the search function in the re library of the Python tool to determine whether the string in the current line exists in the list LS. If it exists, keep it; if not, delete it.

[0075] (5)Match the text of the financial statement notes: Obtain from the knowledge base file the regular expressions start_line = [**Financial Statement Notes|[Company|This Company|Enterprise|Group|This Group] Basic Information] and end_line = [**Supplementary Documents**], where [] represents a set, ** represents other strings, | represents the union, start_line represents the regular expression for the start line of the financial statement notes text, and end_line represents the regular expression for the end line of the financial statement notes text. Use the above regular expressions and the search function of the re library in the python tool to delete the text such as "Important Notice, Table of Contents and Explanation", "Company Profile and Main Financial Indicators" in the financial report that is not part of the financial statement notes. The search function in the re library can be expressed as re.search(), which has three parameters. The first parameter is the matching rule, usually represented by a regular expression. The second parameter is the target string to be matched. The third parameter is the matching mode, with a default value of 0 and can be ignored. When the target string contains the rule represented by the regular expression, re.search() returns an object. When the target string does not contain the rule represented by the regular expression, re.search() returns None. Traverse the TXT document after deleting the header, page numbers, tables, and special symbols. Let the string format of the current line be line. When the return value of re.search(start_line, line) is not None, that is, the expression in start_line is matched in line, start and continue to retain the current line until the return value of re.search(end_line, line) is not None, then stop traversing to obtain the text data of the financial statement notes.

[0076] S2: Identify and label the titles, their levels, and paragraphs in the financial statement notes TXT document, and use the disclosure characteristics of the financial statement notes to assist the event extraction system of the financial statement notes to capture the global information of the document;

[0077] S21: Obtain the TXT document of the financial statement notes text, and obtain the regular expressions for identifying titles and the set of title label styles in the knowledge base as follows:

[0078] (1) Contain Chinese characters, and all Chinese character matches use unicode encoding.

[0079] (2) For those with numbering styles, the numbering style set is [(({0})|({1})|({0})|({1})|{0},|{1},|{0}.|{1}.|{0}.|{1}.|({0}).|({1}).|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|Section {0}|Section {1}|Chapter {0}|Chapter {1})], where [] represents a set, | represents a union, {0} represents the array characters of all Chinese numerals such as "一", "二", etc., and {1} represents the numeric characters of all positive integers such as "1", "2", etc.

[0080] (3) There are no sentence features. The sentence does not include symbols such as [《。 ! ? ;'].

[0081] (4) There are no data features. The sentence does not contain long data, but it can contain years.

[0082] (5) It does not end with a conjunction, that is, the ending does not contain: [and|and|and|or|,|,|”]”|and].

[0083] The lines that meet the above five rules are initially identified as titles and marked with the symbol "S" at the beginning of the line. Rules (1) and (2) are represented as regular expressions and stored in a list, represented by L1. Rules (3), (4), and (5) are represented as regular expressions and stored in a list, represented by L2. Use the search function of the re library in the python tool to determine whether the current line meets the above five rules. Traverse the financial statement notes TXT document and set the string format of the current line to line. When the return value of re.search(L1,line) is not None and the return value of re.search(L2,line) is None, the current line is marked as a title. The marking method is to add the "S" symbol at the beginning of the line and save it as a TXT document.

[0084] S22: Traverse the TXT document of the financial statement notes text, obtain the bigram statistical language model from the model library file, input the line marked with S into the trained bigram statistical language model, and determine whether the current line is a complete semantic sentence. If so, retain the symbol S. If not, delete the symbol S at the beginning of the line to obtain an accurate candidate title.

[0085] The bigram statistical language model in the model library assumes that the current row W is composed of multiple words w1,…,w i ,…,w nComposed, let the probability of W appearing be P(W). According to the conditional probability formula, P(W) can be converted into the product of the conditional probabilities of all words appearing. According to the Markov assumption, the probability of the word w i appearing is only related to w i-1 , then:

[0086]

[0087] Under such an assumption premise, the calculation of P(W) only needs to count the probability of a single word and the simultaneous appearance of the previous and next two words to obtain the probability of W appearing in the current line as P(W). Using the maximum likelihood estimation method, we get:

[0088]

[0089] where count(w i-1 , w i ) represents the word frequency when w i-1 and w i appear simultaneously in the corpus, and count(w i-1 ) represents the word frequency of w i-1 in the corpus.

[0090] In the model library of the present invention, the word frequencies of 33,872 words and the word frequencies of their pairwise combinations have been stored in the form of a list. When judging whether the current line is a sentence with complete semantics, it is necessary to: obtain the accounting term vocabulary in the knowledge base as a word segmentation dictionary, use the accurate mode of the jieba Chinese word segmentation component based on Python to segment the current line W, forming w1,..., w i ,..., w n , search for the corresponding word frequencies count(w1), count(w2),... count(w i ,..., w n ) and the word frequencies of combined phrases count(w1, w2), count(w2, w3),... count(w n ) and count(w n-1 , w n ) in the model library, substitute them into formula (2) to obtain the value of P(w i |w i-1 ), substitute P(w i |w i-1 ) into formula (1) to obtain the value of P(W). When P(W) ≥ 0.8, then retain the symbol S; otherwise, delete the symbol S at the beginning of the line.

[0091] S23: As Figure 3As shown, the TXT document of the financial statement notes is traversed, and the lines marked with S are sorted by numbering style and numbering order. The numbering style set of the title is obtained from the knowledge base. Each numbering style corresponds to a unique code. A numbering hierarchy stack is constructed to record the title level corresponding to each title, and the title is marked with a 10-digit code containing the title level information, including:

[0092] (1) Construct a dictionary TP. The dictionary TP contains multiple key-value pairs. A key-value pair consists of a key and a value connected by a colon, such as {A:" (一)"}, where the key is the code of the label style and the value range is uppercase letters AZ. The value is the label style and the value range is the label style set. There is a unique correspondence between the key and the value. Construct a hierarchical stack to ensure the continuity of the sequence number and the correct hierarchical relationship. Construct a dictionary T to store titles and tag levels.

[0093] (2) Based on the match function of the re library in Python tools, the label style and sequence number of the current line are obtained. The match function in the re library can be expressed as re.match(), which has three parameters. The first parameter represents the matching rule, which is generally expressed as a regular expression. The second parameter represents the target string to be matched. The third parameter is the matching pattern, which defaults to 0 and can be ignored. When the target string contains the rule represented by the regular expression at the beginning, re.match() returns a label style. When the target string does not contain the rule represented by the regular expression, re.match() returns None. The regular expression of the label style is pattern = "symbol + Chinese numerals / Arabic numerals + symbol", where the symbol is "()().、.,,", the symbol can be missing, and the character length of the label is 1-5. Let the string format of the current line be line, use the re.match(pattern,line) function, when re.match(pattern,line) is not equal to None, get the Chinese numerals or Arabic numerals of the current label, convert them into digital format as the serial number n of the current title, when re.match(pattern,line) is equal to None, delete the current title and continue the operation of the next title.

[0094] (3) Use the comparison operator "==" in the Python tool to compare the label style obtained by re.match(pattern,line) with all the label styles in the dictionary TP. If the same style exists, the key corresponding to the return value is the code. If the same style does not exist, determine whether the obtained sequence number n is 1. If the sequence number is 1, record the current label style as the value in the dictionary TP and assign a new key. If the sequence number is not 1, delete the current title and continue the operation of the next title.

[0095] (4) Input the encoding and serial number of the label style into the hierarchical stack. Check if the stack is empty. If it is empty, check if the element to be pushed onto the stack is a title with serial number 1. If it is correct, push it onto the stack. If it is not empty, obtain the top element of the stack and compare the element to be pushed onto the stack with the top element's encoding. Determine the order relationship of the titles based on the encoding and serial number, and push the element to be pushed onto the stack according to the relationship. The relationships are divided into three types: child-level titles, where it is necessary to check the continuity of the encoding and whether the serial number is 1; sibling-level titles, where it is necessary to check the continuity of the serial numbers; parent-level titles, where it is necessary to search for the previous sibling-level title from the top of the stack and determine the continuity of the serial numbers based on the serial number of that title.

[0096] (5) Output the encoding and serial number of the titles in the hierarchical stack in sequence as the values of the keys, and use the content of the titles as the values, and store them in the dictionary T to complete the extraction of the titles and their levels.

[0097] (6) Represent the titles and their levels in the dictionary T using numerical encoding, mark the current title and save it to a TXT document to provide a data basis for paragraph annotation. For example, when the current title is a first-level title and the serial number is 2, mark the current title with 0200000000; when the current title is a third-level title and the serial number is 2, search for the serial number of its parent title. If the second-level parent title is 1 and the first-level parent title is 2, mark the current title with 0201020000.

[0098] S24: Organize the text between titles in the TXT document into one line per paragraph, and mark it with a 10-digit numerical encoding before each paragraph. As Figure 4 shown, obtain the meaning of the encoding, and the title level corresponding to the paragraph is obtained from the annotation of the title. Based on the titles and their levels in step S23, combined with Chinese punctuation marks and serial number symbols, divide the text into paragraphs again, specifically divided into three steps:

[0099] (1) Determine whether the end of a sentence that is not a title ends with [。;!?]. Traverse the TXT document, take the last character of the current line, and use loop statements and the comparison operator "==" in the python tool to compare the last character of the current line with each value in [。;!?] respectively. If there is no equality, keep the current line and continue reading the next line and repeat the operation. If there is equality, it means that the end of the sentence ends with one of the symbols in [。;!?]. Use the logical operation "+" in the python tool to merge all the previously retained lines into one paragraph, add a 10-digit numerical encoding paragraph marker before the paragraph, and then start reading the next line and repeat the operation.

[0100] (2) Based on step (1), determine whether the beginning of the sentence that is not a title is a sequence number style such as ①, (1), (一), 一, (1), (一), 1, etc. If so, merge the paragraphs between the upper and lower sequence number styles into one paragraph. Store the sequence number styles such as ①, (1), (一), 一, (1), (一), 1, etc. in the list LD, traverse the TXT document, take the first five characters of the current line and represent them with line_five, use the re.match(LD, line_five) function to determine whether the beginning of the current line has a sequence number style. When re.match(LD, line_five) is not equal to None, it means that the current line has a sequence number style. Add a paragraph mark in front of the line, continue to read the next line and repeat the operation.

[0101] S3: The semantic features of paragraphs, titles and their levels are represented by vectors, and the vector representations of the words contained in the event arguments and the vector representations of the titles and their levels are concatenated into a vector matrix based on the knowledge of event arguments, providing a data basis for event category classification and event table filling. The specific steps include:

[0102] S31: Input TXT document, match the first 10 digits of each line, divide each document into title set and paragraph set according to the code, use Python tool to intercept the 9th and 10th digits of the current line, and use the comparison operator "==" to determine whether the 9th and 10th digits of the current line are equal to "00". If they are equal, the current line is classified as the title set, if not, the current line is classified as the paragraph set;

[0103] S32: Obtain the trained Doc2evc model from the model library file, obtain the accounting terminology vocabulary from the knowledge base file as the word segmentation dictionary, segment the title set of each document and input it into the Doc2evc model, obtain the vector representation of each row of the title, and obtain the vector list T_V of all the titles;

[0104] A specific example of a vector list T-V is:

[0105] T_V=[[0.67,2.8,0.46,-1.03,1.57,-1.57...-0.34,-0.71,0.78],[0.04,2.11,0.13,0.07,1.2 4-0.24......0.74,-0.7,0.06],......[-0.47,0.44,-0.08,0.15,0.15,0,-0.3,0.31,0.14,0.18]]

[0106] (1) Based on the precise mode of the Python-based jieba Chinese word segmentation component, the current line T i Perform word segmentation and obtain the list [w1,…,wi , …, w n , where w i represents the string of the i-th word.

[0107] (2) Based on 29,280 financial statement note title documents of listed companies, with the help of the gensim.model.doc2vec library in the python tool, a model Doc2evc that integrates upper and lower title information was trained and stored in the model library. The list T i = [w1, …, w i , …, w n was input into Doc2evc, and the vector representation T i V of the current line was obtained by using Doc2evc.infer_vector(T i— ). The length of T i— V is 16. The initialization parameters of Doc2evc are Doc2evc(min_count = 1, window = 5, vector_size = 15, sample = 1e - 3, workers = 4, hs = 1, epochs = 100), where min_count is the minimum frequency of words to be ignored, window is the maximum distance between the current and predicted words in a sentence, vector_size is the dimension of the feature vector, sample is the ratio of random sampling, workers is the number of parallelisms for training, hs sets the optimization solution method (when hs = 1, it is Hierarchical Softmax; when hs = 0, it is Negative sampling;), and epochs is the number of training iterations.

[0108] S33: Obtain accounting term vocabulary from the knowledge base file as a word segmentation dictionary, perform sentence splitting and word segmentation on the paragraph set of each document, and obtain the word segmentation list of the paragraph. Obtain the trained event argument recognition model from the model library file, input the word segmentation list of the paragraph into the event argument recognition model, and obtain the vector matrix P i V that integrates event argument information. The vector matrix includes the feature vectors of all words in the paragraph, and the event arguments are labeled with BIO tags, specifically including:

[0109] (1) Perform sentence splitting and word segmentation on the current paragraph in the accurate mode of the jieba Chinese word segmentation component based on Python, add the SEP tag between sentences, and obtain the word segmentation list P i = [w1, …, w i , …, SEP …, w n , where w i represents the string of the i-th word.

[0110] (2) Input P i into the event argument recognition model to obtain a vector matrix where represents the feature vector of the i-th word. And obtain the labeled tag list P i _L. The tag list is composed of tags of event arguments. For example, P i _L = [o, o, o, …, B-Pledger, I-Pledger, I-Pledger, …, o, o], where o represents a non-event argument, B-Pledger represents the start position of the event argument of the event role Pledger in the sentence, and I-Pledger represents the middle position of the event argument of the event role Pledger in the sentence.

[0111] (3) The event argument recognition model uses the self-attention mechanism in the Transformer encoder to capture the context information of each word in the sentence, and then classifies through the CRF model to predict the label of each word. The Transformer encoder is composed of 6 encoding blocks. After the input data passes through the self-attention mechanism module, a weighted feature vector is obtained where Q, K, and V are obtained by linearly transforming the input data P i , d k is the number of columns of the Q and K matrices, that is, the vector dimension. First, convert the word segmentation list P i into an embedding vector, and then obtain Q = W Q P i , K = W K P i , V = W V P i through linear transformation. W Q , W K and W V are trainable parameter matrices. Before model training, W Q , W K and W V are randomly initialized.

[0112] (4) Input Z into the feed-forward neural network layer (FFN) to output the feature vector FFN(Z) in the compressed space. FFN has two layers. The activation function of the first layer is ReLU, and the second layer is a linear activation function, expressed as: FFN(Z) = max(0, ZW1 + b1)W2 + b2. The output of the last feed-forward neural network layer is the vector matrix Input P i _V into the CRF model as input to obtain the conditional probability that the output label of each sample w i is y: where t k 、s l : Feature function; λ k 、u l : Corresponding weight value; The feature function is a predefined rule function, and λ k 、u l are parameters obtained during the training of the model.

[0113] (5) Obtain the conditional probability value P(y / x) according to the output of the CRF model, and output the vector matrix P i _V corresponding annotation label list P i _L.

[0114] Label list:

[0115] P i _L = [o, B-SigningTime, I-SigningTime, I-SigningTime, o…, B-Terms, I-Terms, I-Terms, I-Terms, o]

[0116] S34: Represent the title and its hierarchical vector T i— V and the vector representation P of the words included in the event argument i _V are concatenated into a vector matrix PE i , and only the word vectors with the label being the event argument are retained in the vector matrix PE i , and the first 16 bits of the word vectors in the same paragraph are the feature information of the same title and its hierarchy. When the same paragraph corresponds to multiple-level titles, the title vectors of all levels are averaged and then concatenated. That is where represents the average value of the title vectors of all levels corresponding to the word w i , and represents the word vector corresponding to the word w i .

[0117] S4: Construct a Transformer encoder and a softmax classifier, and judge the event category by learning the features of the event argument and the title and its hierarchy.

[0118] S41: Input the vector matrix PE that fuses the event argument and the title and its hierarchy i into the Transformer encoder for feature learning, which is the same as the Transformer encoder in the event argument recognition model. Input PE i , and after passing through the encoding block composed of 6 self-attention mechanisms and feed-forward neural network layers, output the vector matrix E that fuses the event argument and the title and its hierarchy information = [e1, e2…e i ,…, en , i.e., the vector matrix of enhanced event arguments, where e i represents the vector representation of the i-th event argument after integrating the title and its hierarchical information.

[0119] S42: Connect the output layer of the Transformer encoder to the event classifier for event category classification, obtain the event category triggered by the current paragraph, and obtain the label of the event category triggered by the current paragraph by taking the maximum probability, which is specifically expressed as:

[0120] max(softmax(WE + b)) = max([0, 0.03, 0, …, 0.05, 0.89]) = 0.89

[0121] label(31) = "S32 Major Contract"

[0122] where the event classifier is composed of a softmax classifier. The specific operation is to input E = [e1, e2, …, e i , …, e n into the softmax classifier and calculate where and are learnable parameter matrices. Here, k refers to the number of event category labels. 32 event categories are predefined, so k = 32. And is the output of the softmax classifier, which is a k-dimensional vector containing the confidence rates of all event categories.

[0123] Use cross-entropy to represent the information difference between the label yi of the true event type and the result predicted by the model. Then the optimization objective is to make the cross-entropy as small as possible. Define the loss function as the average cross-entropy where SE is the total number of event types, pic is the predicted probability of type c in . Randomly initialize W and b, and use gradient descent to update the optimization variables.

[0124] S43: Obtain the file of the event category from the knowledge base file, obtain the composition of the predefined event roles in the current event category, and form the column names of the event table in the order of the predefined event roles.

[0125] The event category and event role can be customized according to the requirements of the extraction task. For example, the event role list of the current event category is expressed as:

[0126] Major Contracts: ["TransactionSubject", "TransactionObject", "SigningTime", "Terms", "ContractAmount", "State", "ProjectDuration"]

[0127] S5: Construct a Transformer encoder and a linear binary classifier. By learning the features of event arguments, titles, their hierarchies, and memory vectors, determine whether each event argument is filled into the current event role. The specific steps are as follows:

[0128] S51: Traverse the list of event roles, initialize the vector representation ER of the event role and the memory vector M, select the first event role, vectorize the label of the event role, concatenate it with the vectorized representations E of all event arguments in the current paragraph, and at the same time concatenate the memory vector M into the vector representation E of the event argument to obtain a vector matrix

[0129] S52: Input the vector matrix into the Transformer encoder, which is the same as the Transformer encoder in the event argument recognition model. After passing through an encoding block composed of 6 self-attention mechanisms and feed-forward neural network layers, output a vector matrix that fuses the information of event arguments, titles, their hierarchies, and memory vectors Input into the event table filling unit to obtain the probabilities [0, 0, 0, 1, 0,......, 0] that all event arguments can be filled into the current event role.

[0130] S53: Connect the output layer of the Transformer encoder to the linear binary classifier to determine whether all event arguments are filled in this event role. The linear binary classifier is calculated using the logistic regression formula, that is p i represents the probability that the i-th argument can be filled into the current event role, and W and b are learnable parameter matrices. During model training, its loss function is Randomly initialize W and b, and use gradient descent to update and optimize the variables.

[0131] S54: Update the memory vectors of the filled event arguments and the vectorized representations of the event roles.

[0132] S55: Repeat steps S52, S53, and S54 until all event roles of the current event category are filled. Event roles without event arguments are filled with NULL values to obtain all event records contained in the current paragraph.

[0133] In this embodiment, the output layer of the Transformer encoder is connected to the CRF model to construct an event argument recognition model, aiming to complete the event argument recognition task by using the feature extraction function of the Transformer encoder and the classification function of the CRF model. Previous event extraction models cannot capture the discourse semantic information of the document when applied to the text of financial statement footnotes. The present invention mainly adds the information of the title and its hierarchy as the discourse semantic information of the document before the event argument recognition model, improving the accuracy of financial statement footnote event classification, and learning the features of event arguments, titles and their hierarchies, and memory vectors through the Transformer encoder, where the memory vector records the classification results of event arguments. On this basis, a linear binary classifier is added to judge whether the event argument matches the event role. The Transformer encoder and the linear binary classifier are used to transform the event table filling problem into a binary classification problem of whether the event argument matches the target event role, solving the problem of having multiple event records for the same event category and further improving the accuracy of financial statement footnote event extraction.

[0134] Embodiment 2

[0135] As Figure 5 shown, this embodiment provides an event extraction system for financial statement footnotes, including: a document acquisition unit, a discourse structure recognition unit, an event argument recognition unit, an event category classification unit, and an event table filling unit;

[0136] The document acquisition unit, as the interface of the event extraction system of the present invention, is used to preprocess the document that needs to implement the event extraction task, acquire the PDF document of the financial report in the database file, convert the PDF document into a TXT document after data preprocessing, and delete the text that is not the text of the financial statement footnotes;

[0137] Based on the document acquisition unit, the discourse structure recognition unit obtains a title set and a paragraph set by identifying and annotating the titles and their hierarchies and paragraphs in the financial statement footnote TXT document. The titles and their hierarchies in the title set are represented by vectors using the Doc2evc model and input into the event category classification unit. This module is to assist the event extraction system of the present invention in capturing the discourse semantic information of the document;

[0138] The event argument recognition unit represents the semantic features of the paragraph into vectors through the learning task of event argument recognition. First, the paragraph set output by the discourse structure recognition unit is segmented into sentences and words to obtain a word segmentation list, then the Transformer encoder is used to learn the semantics in the paragraph, and then the vector matrix of the output layer of the Transformer encoder is input into the CRF model to identify and annotate the event arguments of the financial events in the financial statement footnotes, obtaining the vector representation of the event arguments.

[0139] The event category classification unit designs a vector splicing mechanism, which splices the vector representations of the words contained in the event arguments and the vector representations of the title and its hierarchy into a vector matrix, providing a data basis for event category classification and event table filling. The spliced vector matrix is input into the Transformer encoder to obtain a vector matrix that fuses the title, title hierarchy, and paragraph semantics. This vector matrix is then input into the event category classifier to obtain the probabilities of all event categories for the current vector matrix. The event category with the highest probability is selected as the event category of the current vector matrix. The event roles of the currently triggered event category in the predefined event information are queried through indexing, and the event roles are output in the predefined order.

[0140] The event table filling unit is based on the Transformer encoder and a linear binary classifier. The input is the vector representation e of the event arguments output by the event argument recognition unit and the vector representation of the event roles output by the event category classification unit. To enable an event role to fill more than one event argument with good accuracy, the system adds a memory vector m to record the filling process of the event arguments. After splicing the vectors of the event arguments, memory vector, and event roles, they are input into the Transformer encoder. Then, the output layer of the Transformer encoder is connected to the linear binary classifier to obtain the probability that the event argument fills the current event role. The event argument with a probability of 1 is selected to fill the current event role in the event table. The above operations are looped in the predefined order of event roles until the event role filling is completed.

[0141] Embodiment 3

[0142] This embodiment provides a computer-readable storage medium, which can be a storage medium such as ROM, RAM, disk, or optical disc. The storage medium stores one or more programs. When the programs are executed by a processor, the event extraction method for the financial statement notes in Embodiment 1 is implemented.

[0143] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for event extraction of the notes to financial statements, characterized in that, It includes the following steps: Obtain the PDF document of the financial report in the database file. After data preprocessing, convert the PDF document into a TXT document, and combine the regular expressions in the knowledge base to match the TXT document of the financial statement notes text; Based on the knowledge base, identify and label the titles, title levels, and paragraphs of the TXT document of the financial statement notes text to obtain a title set and a paragraph set; Perform sentence splitting and word segmentation on the paragraph set to obtain a word segmentation list. Based on the Transformer encoder, learn the semantics in the paragraphs, and input the vector matrix of the output layer of the Transformer encoder into the CRF model to identify and label the event arguments of the financial events in the financial statement notes to obtain the vector representation of the event arguments; Concatenate the vector representations of the words included in the event arguments and the vector representations of the titles and their levels into a vector matrix. Input the concatenated vector matrix into the Transformer encoder to obtain a vector matrix that fuses the semantics of the titles, title levels, and paragraphs. Input the vector matrix into the event category classifier to obtain the probabilities of all event categories of the current vector matrix. Select the event category with the highest probability as the event category of the current vector matrix. Query the event roles of the currently triggered event category in the predefined event information through indexing, and output the event roles in the predefined order to obtain the vector representation of the event roles; Construct a memory vector for recording the event argument filling process. Concatenate the vector representation of the event arguments, the vector representation of the event roles, and the memory vector and input them into the Transformer encoder. Connect the output layer of the Transformer encoder to a linear binary classifier to obtain the probability of the event argument filling the current event role. Select the event argument with a probability equal to the set value and fill it into the current event role of the event table. Repeat the iteration until all event roles of the current event category are filled, and update the memory vector of the filled event arguments and the vectorized representation of the event roles to obtain all event records included in the current paragraph.

2. The method for event extraction of the notes to the financial statements according to claim 1, wherein Data preprocessing includes deleting headers, page numbers, tables, and special symbols.

3. The method for event extraction of the notes to the financial statements according to claim 1, characterized in that Combine the regular expressions in the knowledge base file to match the TXT document of the financial statement notes text. The regular expression is expressed as: start_line = [**Financial Statement Notes|Basic Information of [Company|This Company|Enterprise|Group|This Group]] and end_line = [**Supplementary Documents**]; Among them, [] represents a set of regular expressions, ** represents other strings, | represents an optional item, start_line represents the regular expression of the start line of the financial statement notes text, and end_line represents the regular expression of the end line of the financial statement notes text; When the target string contains the rules represented by the regular expression, re.search() returns an object. When the target string does not contain the rules represented by the regular expression, re.search() returns None; Traverse the TXT document after data preprocessing, set the string format of the current line to line, when the return value of re.search(start_line, line) is not None, that is, the expression in start_line is matched in line, start and continue to retain the current line until the return value of re.search(end_line, line) is not None, then stop traversal and obtain the TXT document of the financial statement notes text.

4. The method for event extraction of the notes to the financial statements according to claim 1, wherein Identify and annotate the titles, title levels and paragraphs of the TXT documents of the notes to the financial statements, including: Obtain the TXT document of the financial statement notes, obtain the regular expression for identifying the title and the label style set of the title in the knowledge base, and add a mark symbol at the beginning of the line identified as the title; Traverse the TXT document of the financial statement notes text, and determine whether the line with the added marker is a complete semantic sentence based on the binary statistical language model. If it is a complete semantic sentence, keep the marker, otherwise delete the marker at the beginning of the line to obtain the candidate title; Traverse the TXT document of the financial statement notes, sort out the numbering styles and numbering order of the lines with added mark symbols, load the numbering style set of the title from the knowledge base, each numbering style corresponds to a unique code, build a numbering hierarchy stack to record the title level corresponding to each title, and annotate the title with a digital code containing the title level information; The text between titles in the TXT document is organized into a line according to paragraphs, and a digital code is used before the paragraph to obtain the meaning of the code. The title level corresponding to the paragraph is obtained from the title mark. Based on the title and its level, the text is divided into paragraphs again in combination with Chinese punctuation marks and serial number symbols.

5. The method for event extraction from the notes to the financial statements according to claim 4, wherein Get the regular expression for identifying titles and the title labeling style set in the knowledge base. The specific rules are expressed as follows: The first rule: contains Chinese characters, and Chinese characters are matched using unicode encoding; Second rule: For those with label styles, the label style set is [(({0})|({1})|({0})|({1})|{0},|{1},|{0}.|{1}.|{0}.|{1}.|({0}).|({1}).|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|({0}),|({1}),|Section {0}|Section {1}|Chapter {0}|Chapter {1})], where [] represents a set, | represents an optional item, {0} represents the array characters of all Chinese numerals, and {1} represents the numeric characters of all positive integers; The third rule: If there is no sentence feature, the sentence does not include the [《。 ! ? ;'] symbol; The fourth rule: no data features, no long data in the sentence, but can contain the year; Rule 5: Do not end your sentence with a conjunction; Lines that satisfy the first to fifth rules are initially identified as titles. The first and second rules are represented as regular expressions and stored in a list denoted as L1. The third, fourth, and fifth rules are represented as regular expressions and stored in a list denoted as L2. Traverse the TXT document of the financial statement notes. Let the string format of the current line be line. When the return value of re.search(L1, line) is not None and the return value of re.search(L2, line) is None, the current line is marked as a title. re.search is the search function in the re library of the Python tool.

6. The method for event extraction of the notes to the financial statements according to claim 4, wherein Construct a hierarchical stack of labels to record the title level corresponding to each title, and label the title with a numeric code containing title level information. Specifically, it includes: Construct a dictionary TP. The dictionary TP contains multiple key-value pairs. A key-value pair consists of a key and a value, connected by a colon. Among them, the key is the code of the label style, and the value is the label style. There is a unique correspondence between the key and the value. Construct a dictionary T to store the title and its marked level; Obtain the label style and serial number of the current line. Compare the obtained label style with all the label styles in the dictionary TP. If there is a similar style, return the corresponding key, that is, the code. If there is no similar style, judge whether the obtained serial number n is 1. If the serial number is 1, record the current label style as the value in the dictionary TP and assign a new key. If the serial number is not 1, delete the current title; Input the code and serial number of the label style into the hierarchical stack. Judge whether the stack is empty. If it is empty, judge whether the element to be pushed onto the stack is a title with a serial number of 1. If it is correct, push it onto the stack. If it is not empty, obtain the top element of the stack and compare the element to be pushed onto the stack with the top element of the stack. Judge the sequence relationship of the titles according to the code and serial number, and push the element to be pushed onto the stack according to the relationship between the titles; Output the codes and serial numbers of the titles in the hierarchical stack in sequence as the value of the key, and use the content of the title as the value, and store them in the dictionary T to complete the extraction of the title and its level; Represent the title and its level in the dictionary T with numeric codes, and mark the current title and save it to the TXT document.

7. The method for event extraction of the notes to the financial statements according to claim 1, characterized in that, Perform sentence splitting and word segmentation on the paragraph set to obtain a word segmentation list. Specifically, it includes: Tokenize the current line T i to obtain a list [w1, …, w i , …, w n , where w i represents the string of the i-th word. Input the list T i = [w1, …, w i , …, w n into the Doc2evc model to obtain the vector representation T i —V; Perform clause splitting and word segmentation on the current paragraph, add the SEP tag between sentences, and obtain the word segmentation list P i = [w1,…,w i ,…,SEP…,w n ; Learn the semantics in the paragraph based on the Transformer encoder. Specifically, use the self-attention mechanism of the Transformer encoder to capture the context information of each word in the sentence. After passing through the self-attention mechanism module, the input word segmentation list obtains a weighted feature vector to obtain a vector matrix; Obtain the conditional probability value of each sample output as the corresponding label according to the CRF model, and output the annotation label list corresponding to the vector matrix.

8. The method for event extraction from the notes to the financial statements according to claim 1, characterized in that, Input the spliced vector matrix into the Transformer encoder to obtain a vector matrix that fuses the title, title hierarchy, and paragraph semantics, denoted as: E = [e1, e2…e i ,…,e n , where e i represents the vector representation of the i-th event argument after integrating the title and its hierarchical information; Connect the output layer of the Transformer encoder to the event classifier for classifying event categories, obtain the event category triggered by the current paragraph, and obtain the label of the event category triggered by the current paragraph by taking the maximum probability. Input the vector matrix E = [e1, e2…e i ,…,e n into the softmax classifier and calculate where W and b represent learnable parameter matrices, and k represents the number of event category labels; Use cross-entropy to represent the information difference between the label of the true event type and the prediction result output by the softmax classifier. Define the loss function as the average cross-entropy, randomly initialize W and b, and update using gradient descent.

9. An event extraction system for the notes to financial statements, characterized in that, Include: A document acquisition unit, a discourse structure recognition unit, an event argument recognition unit, an event category classification unit, and an event table filling unit; The document acquisition unit is used to acquire the PDF document of the financial report in the database file. After data preprocessing, the PDF document is converted into a TXT document, and the TXT document of the financial statement notes text is matched by combining the regular expressions in the knowledge base. The discourse structure recognition unit is used to identify and label the titles, their levels and paragraphs of the TXT document of the financial statement notes text based on the knowledge base, obtain a title set and a paragraph set, represent the titles and their levels in the title set as vectors based on the Doc2evc model, and input them into the event category classification unit. The event argument recognition unit is used to split the paragraphs in the paragraph set into sentences and words to obtain a word segmentation list, learn the semantics in the paragraphs based on the Transformer encoder, input the vector matrix of the output layer of the Transformer encoder into the CRF model, identify and label the event arguments of the financial events in the financial statement notes, and obtain the vector representation of the event arguments. The event category classification unit is used to splice the vector representations of the words included in the event arguments and the vector representations of the titles and their levels into a vector matrix, input the spliced vector matrix into the Transformer encoder to obtain a vector matrix that fuses the title, title hierarchy and paragraph semantics, input the vector matrix into the event category classifier to obtain the probabilities of all event categories of the current vector matrix, select the event category with the highest probability as the event category of the current vector matrix, query the event roles of the currently triggered event category in the predefined event information through indexing, and output the event roles in the predefined order to obtain the vector representation of the event roles. The event table filling unit is used to construct a memory vector for recording the event argument filling process, splice the vector representation of the event arguments, the vector representation of the event roles and the memory vector and input them into the Transformer encoder, connect the output layer of the Transformer encoder to a linear binary classifier to obtain the probability that the event argument fills the current event role, select the event argument with the probability being the set value and fill it into the current event role of the event table, repeat the iteration until all event roles of the current event category are filled, update the memory vector of the filled event arguments and the vectorized representation of the event roles, and obtain all event records included in the current paragraph.

10. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the event extraction method for the financial statement notes as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Document-level event argument extraction method based on sequence labeling

    CN113591483A

  • System and methods for automated detection, reasoning and recommendations for resilient cyber systems

    US20180103052A1