A method and device for extracting contract information
By identifying and analyzing the article title and its hierarchy and key text information of the contract file, combining long text pre-training models and regular expressions, the problem of poor format adaptability in the prior art is solved, and efficient and accurate contract information extraction is achieved.
Patent Information
- Application Number
- CN202111438732.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-26
AI Technical Summary
The prior art requires writing a large number of rules to adapt to different formats when extracting document-level information, with poor generalization, unable to filter interference information, and unable to extract specific key indicator information.
By identifying the article title and its hierarchy and key text information, combining text, layout and visual features, long text pre-trained models and regular expressions are used to automatically extract key information from the contract file.
It realizes efficient and accurate extraction of contract information in documents in different formats, improves the universality and accuracy of information extraction, and reduces dependence on rule writing.
Smart Images

Figure CN114118053B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of information extraction, and particularly to a method and device for extracting contract information. Background Art
[0002] In the current field of information extraction, when performing document-level information extraction, the main current idea is to first parse the PDF in combination with OCR technology, retain information such as its text and location, then use document distribution rules combined with deep learning to identify its title and restore the document layout, and finally extract key information based on the title and text keywords.
[0003] However, traditional technical means require writing a large number of rules to fit various PDF formats, and have poor generalization, with a poor fitting effect for new types of layouts, and cannot filter out interference information such as tables of contents and multiple documents, nor can they extract specific key index information.
[0004] Therefore, to meet the usage requirements, a contract information extraction technology is provided. Summary of the Invention
[0005] This application provides a method and device for extracting contract information, which does not require writing rules according to different layouts, and extracts contract information by identifying the article title and its corresponding title level, combined with the corresponding key text information, improving the generality of technical implementation while ensuring the accuracy of information extraction.
[0006] In a first aspect, this application provides a method for extracting contract information, and the method includes the following steps:
[0007] Receive a contract file, perform text parsing, and obtain text parsing data;
[0008] Based on the text parsing data, obtain multiple contract text paragraphs of the contract file;
[0009] Based on the text feature information corresponding to the contract text paragraphs, obtain article titles at different levels of the contract file;
[0010] Based on the text feature information corresponding to the contract text paragraphs corresponding to the article titles, identify and obtain the key text information corresponding to the article titles; where
[0011] The text parsing data includes:
[0012] Text feature information, which is used to record the text content in the contract file;
[0013] Layout feature information, which is used to record the orientation corresponding to the text content;
[0014] Visual feature information, which is used to record the font and font size corresponding to the text content.
[0015] Specifically, the text feature information includes multiple text feature sub-information, and the text feature sub-information is used to record any paragraph of text content in the contract file;
[0016] The text feature sub-information respectively corresponds to the article title at any level.
[0017] Specifically, in the process of identifying the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0018] Based on the contract text paragraph corresponding to the article title, obtain the text feature sub-information corresponding to the corresponding contract text paragraph;
[0019] Identify the text feature sub-information to identify the key text information corresponding to the article title; wherein,
[0020] The text feature information includes multiple text feature sub-information;
[0021] The text feature sub-information corresponds to a contract text paragraph.
[0022] Specifically, in the process of identifying the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0023] Based on the contract text paragraph corresponding to the article title, obtain the text feature sub-information corresponding to the corresponding contract text paragraph;
[0024] Identify the text feature sub-information and compare it with the case sentences of different case texts in the preset case library to compare the similarity with the case text;
[0025] Based on the case text with the best similarity, obtain the key text information corresponding to the article title; wherein,
[0026] The text feature information includes multiple text feature sub-information;
[0027] The text feature sub-information corresponds to a contract text paragraph.
[0028] Specifically, in the process of obtaining the article titles at different levels of the contract file based on the text feature information corresponding to the contract text paragraph, the following steps are included:
[0029] Obtain the article title in the contract document from the text feature information based on the layout feature information and the visual feature information;
[0030] Based on the layout feature information and the visual feature information, identify the hierarchical relationship between different article titles.
[0031] Further, the method further includes the following steps:
[0032] Based on each article title, the hierarchical relationship between different article titles, and the key text information corresponding to each article title, establish a corresponding association relationship.
[0033] In a second aspect, the present application provides a contract information extraction device, and the device includes:
[0034] A file parsing module, which is used to receive a contract document, perform text parsing, and obtain text parsing data;
[0035] A layout restoration module, which is used to obtain multiple contract text paragraphs of the contract document based on the text parsing data, and is also used to obtain different levels of article titles of the contract document based on the text feature information corresponding to the contract text paragraphs;
[0036] An information extraction module, which is used to identify and obtain the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title; where
[0037] The text parsing data includes:
[0038] Text feature information, which is used to record the text content in the contract document;
[0039] Layout feature information, which is used to record the orientation corresponding to the text content;
[0040] Visual feature information, which is used to record the font and font size corresponding to the text content;
[0041] The text feature information includes multiple text feature sub-informations, and the text feature sub-informations are used to record any paragraph of text content in the contract document;
[0042] The text feature sub-informations respectively correspond to any level of the article title.
[0043] Further, the information extraction module is further used to obtain the text feature sub-information corresponding to the contract text paragraph corresponding to the article title based on the contract text paragraph corresponding to the article title, and is also used to identify the text feature sub-information, and identify and obtain the key text information corresponding to the article title; where
[0044] The text feature information includes multiple text feature sub-informations;
[0045] The text feature sub-information corresponds to a contract text paragraph.
[0046] Furthermore, the information extraction module is also used to obtain the corresponding text feature sub-information of the contract text paragraph based on the contract text paragraph corresponding to the article title;
[0047] The information extraction module is also used to identify the text feature sub-information, compare it with the case sentences of different case texts in the preset case library, and compare the similarity with the case text;
[0048] The information extraction module is also used to obtain the key text information corresponding to the article title based on the case text with the best similarity; wherein,
[0049] The text feature information includes multiple text feature sub-informations;
[0050] The text feature sub-information corresponds to a contract text paragraph.
[0051] Furthermore, the layout restoration module is also used to obtain the article title in the contract document from the text feature information based on the layout feature information and the visual feature information;
[0052] The layout restoration module is also used to identify the hierarchical relationship between different article titles based on the layout feature information and the visual feature information.
[0053] The beneficial effects brought by the technical solution provided by this application include:
[0054] This application does not need to write rules according to different layouts. By identifying the article title and the corresponding title hierarchy, and combining the corresponding key text information, contract information extraction is carried out, which improves the generality of technical implementation while ensuring the accuracy of information extraction. Brief Description of the Drawings
[0055] Term Explanation:
[0056] pdf: Portable Document Format, a portable document format;
[0057] OCR: Optical Character Recognition, optical character recognition;
[0058] MLM: Masked Language Model, a masked language model.
[0059] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0060] Figure 1 It is a step flowchart of the contract information extraction method provided in the embodiments of the present application;
[0061] Figure 2 It is a schematic diagram of the contract information extraction method provided in the embodiments of the present application;
[0062] Figure 3 It is a structural block diagram of the contract information extraction device provided in the embodiments of the present application. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0064] The following further elaborates on the embodiments of the present application with reference to the accompanying drawings.
[0065] The embodiments of the present application provide a contract information extraction method and device, which do not need to write rules according to different formats. By identifying the article title and the corresponding title level, combined with the corresponding key text information, contract information extraction is performed, and the generality of technical implementation is improved on the premise of ensuring the accuracy of information extraction.
[0066] To achieve the above technical effects, the general idea of the present application is as follows:
[0067] A contract information extraction method, the method includes the following steps:
[0068] S1. Receive a contract file, perform text parsing, and obtain text parsing data;
[0069] S2. Based on the text parsing data, obtain multiple contract text paragraphs of the contract file;
[0070] S3. Based on the text feature information corresponding to the contract text paragraphs, obtain article titles at different levels of the contract file;
[0071] S4. Based on the text feature information corresponding to the contract text paragraph corresponding to the article title, identify and obtain the key text information corresponding to the article title; among them,
[0072] The text parsing data includes:
[0073] Text feature information, which is used to record the text content in the contract document;
[0074] Layout feature information, which is used to record the orientation corresponding to the text content;
[0075] Visual feature information, which is used to record the font and font size corresponding to the text content.
[0076] The following further elaborates on the embodiments of the present application with reference to the accompanying drawings.
[0077] In the first aspect, as shown in Figures 1 - 2 the embodiments of the present application provide a contract information extraction method, which includes the following steps:
[0078] S1. Receive a contract document, perform text parsing, and obtain text parsing data;
[0079] S2. Based on the text parsing data, obtain multiple contract text paragraphs of the contract document;
[0080] S3. Based on the text feature information corresponding to the contract text paragraph, obtain different levels of article titles of the contract document;
[0081] S4. Based on the text feature information corresponding to the contract text paragraph corresponding to the article title, identify and obtain the key text information corresponding to the article title; among them,
[0082] The text parsing data includes:
[0083] Text feature information, which is used to record the text content in the contract document;
[0084] Layout feature information, which is used to record the orientation corresponding to the text content;
[0085] Visual feature information, which is used to record the font and font size corresponding to the text content.
[0086] Specifically, the text feature information includes multiple text feature sub-information, and the text feature sub-information is used to record any paragraph of text content in the contract document;
[0087] The text feature sub-information respectively corresponds to any level of the article title.
[0088] It should be noted that the text feature information corresponds to the literal content of the contract document, and the contract document can be specifically divided into literal contents of multiple paragraphs, denoted as contract text paragraphs. Therefore, when parsing, the literal content of the contract document can also be obtained paragraph by paragraph, denoted as different text feature sub-information. One text feature sub-information corresponds to one contract text paragraph, and each paragraph corresponds to article titles at different levels. Therefore, each of the text feature information corresponds to an article title, that is, each of the text feature information belongs to an article title corresponding to it;
[0089] In addition, due to the level characteristics of article titles at different levels, at least one contract text paragraph, that is, at least one text feature sub-information, is included under one article title;
[0090] When the content of the article title is too much, at least two or even multiple contract text paragraphs can be included under one article title, that is, two or even multiple text feature sub-information can be included.
[0091] Among them, in step S1, a pdf parsing module can be trained using a contract document dataset, specifically, an OCR recognition module can be trained based on the open-source project pdfminer combined with an OCR model;
[0092] For the input pdf file, the output is a contract parsing text, that is, text parsing data, and its content includes at least text information, coordinates for indicating the position of the text, font, and font size.
[0093] In the embodiment of the present application, there is no need to write rules according to different layouts, and by identifying the article title and the corresponding title level, combined with the corresponding key text information, to extract contract information, which improves the generality of technical implementation while ensuring the accuracy of information extraction.
[0094] It should be noted that in the embodiment of the present application, the contract document can specifically be a file in pdf format.
[0095] Specifically, in step S3, in obtaining different-level article titles of the contract document based on the text feature information corresponding to the contract text paragraph, the following steps are included:
[0096] Based on the layout feature information and the visual feature information, obtain the article title in the contract document from the text feature information;
[0097] Based on the layout feature information and the visual feature information, identify the hierarchical relationship between different article titles.
[0098] Based on the technical solution of the embodiment of the present application, in step S2, the paragraphs in the contract document can be split according to the text parsing data, so as to obtain multiple contract text paragraphs of the contract document;
[0099] The data basis for paragraph splitting can be the text feature information, layout feature information, and visual feature information of the text parsing data. When necessary, the corresponding punctuation marks can also be combined.
[0100] It should be noted that step S3, in specific implementation, includes at least two stages, and the specific situation is as follows:
[0101] The first stage is to identify the article title in the document:
[0102] The training uses the content of each page of the pdf document as a sample unit, and adopts the longformer pre-trained model. The pre-trained base model is the publicly available Chinese model longformer-chinese-base-4096, which is optimized and trained on Roberta. By optimizing the self-attention structure of the transformer, the computational complexity is reduced, so that the model can model long texts;
[0103] When necessary, the open-source pre-trained model longformer-chinese-base-4096 can also be retrained. The training corpus is the contract text in the industry, and the training method is MLM. Through further masked training, the model is more sensitive to the vertical industry.
[0104] The inputs are text features, layout features, and visual features. The text features are mainly the text content in the pdf, which is input sentence by sentence. The layout features are the coordinate box information x0, y0, x1, y1 corresponding to the text box, which respectively correspond to the coordinates of the upper left corner and the lower right corner of the text box, and the height and width of each character are calculated through the coordinates. The visual features are information such as the font and font size of the text;
[0105] Finally, there are 9 features including token, x0, y0, x1, y1, width, height, fontname, and fontsize as inputs. The lengths of each feature are unified to 2048, and padding is performed for those that are insufficient. Among them, relevant calculations on the token can generate position and segment features. The three feature vectors of token, position, and segment are used as the input of longformer, and a vector embedding1 with dimensions of B*T*E will be obtained;
[0106] Round the coordinate data of x0, y0, x1, y1, width, and height. The index range is 0 - 1024. Obtain embedding2 through vector embedding with random initialization; Use all the data to construct a dictionary for indexing fontname and fontsize, and obtain embedding3 through vector embedding with random initialization. Perform weighted summation on the three output vectors, and then connect to a fully connected layer to construct a binary classifier to achieve article title recognition.
[0107] It should be noted that the contract file in pdf format includes multiple text boxes, which together constitute the text content of the pdf file.
[0108] Second stage: Identify the hierarchy of titles in the document
[0109] In this stage, the longformer pre-trained model can also be used. The input data for this stage is the output data of the first stage. The difference is that the input of the first stage is all text, while the input of the second stage is the article title. The training uses the title of each pdf document as the sample unit, and the processing methods of features are the same, and the intermediate layer network structures are the same;
[0110] During the process of identifying titles and their hierarchies, relevant rules will be added to correct the model results. After the first stage of recognition is completed, construct a set of title sentences according to the corresponding font sizes, and then further search for title lines with obvious identifiers in the text through regular expressions, such as sentences starting with "the first" or "1.1". After matching relevant sentences, it will be checked whether they are in the set generated by the model. If they are, the sentence will be output as the title in the end; if not, it will not be output.
[0111] In the second stage, a similar method is also constructed. Build a font size dictionary from the output of the model. Each level of title corresponds to a unique font size. Use regular expressions to search all titles. Centered titles are usually the first-level titles. In this way, part of the model results will be corrected. Finally, a hierarchical number like 1, 2, 2, 2, 3, 1, 1, 2 will be generated, and then it will be converted into a table of contents number with a structure of 1, 1.1, 1.2, 1.3, 1.3.1, 2, 3, 3.1. In addition, title identifiers like 1.19 and 2.3 usually correspond to the 19th and 3rd titles under the corresponding title level in the end. The title numbers under the level can be corrected through this method;
[0112] Finally, the corrected results are stored in a dictionary structure of key-value. The key is the title and its corresponding level, and the value is the text content under the title.
[0113] It should be noted that in the second stage, the pipeline structure is used to identify the document title in stages. This method avoids writing a large number of rules, has strong migration ability, is not restricted by rules, and can greatly improve the accuracy and effect of the model in combination with manual rule correction.
[0114] Specifically, in step S4, in the process of identifying the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0115] Based on the contract text paragraph corresponding to the article title, obtain the corresponding text feature sub-information of the contract text paragraph;
[0116] Identify the text feature sub-information to identify the key text information corresponding to the article title; among them,
[0117] The text feature information includes multiple text feature sub-informations;
[0118] The text feature sub-information corresponds to a contract text paragraph.
[0119] Specifically, there is an implementation case of step S4, that is, in the process of identifying the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0120] Based on the contract text paragraph corresponding to the article title, obtain the corresponding text feature sub-information of the contract text paragraph;
[0121] Identify the text feature sub-information to identify the key text information corresponding to the article title; among them,
[0122] The text feature information includes multiple text feature sub-informations;
[0123] The text feature sub-information corresponds to a contract text paragraph, which corresponds to the article title at any level.
[0124] Specifically, there is another implementation case of step S4, that is, in the process of identifying the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0125] Based on the contract text paragraph corresponding to the article title, obtain the corresponding text feature sub-information of the contract text paragraph;
[0126] Identify the text feature sub-information, compare it with the case sentences of different case texts in the preset case library, and compare the similarity with the case text;
[0127] Obtain the key text information corresponding to the article title based on the best-case text in terms of similarity; wherein,
[0128] The text feature information includes multiple text feature sub-information;
[0129] The text feature sub-information corresponds to a paragraph of the contract text;
[0130] The text feature sub-information corresponds to the article title at any level.
[0131] It should be noted that in specific implementation, keyword fields will be preset, and corresponding screening rules will be configured. In the way of integrating rules and algorithms, if relevant rule methods such as regular expressions and artificial logics fail to recognize, further matching and retrieval will be performed in the case library. Step S4 includes the following two working conditions:
[0132] Condition 1: The information corresponding to the keyword field will be directly set under the corresponding subtitle in some documents. For example, the confidentiality clause information. In this case, there will be a special paragraph to elaborate on it, and its title is the confidentiality clause. For this category, the title can be directly indexed in the dictionary and output.
[0133] Condition 2: The information corresponding to the keyword field will be hidden in the text. For example, Party A: xxxxxx Company, Project Name: xxxxxx. For this situation, regular expressions are used for extraction and output;
[0134] For Condition 2, in specific implementation, several sentences corresponding to this field will be input first, such as 6 - 10 sentences, and stored in the case library as cold-start matching sentences. In the formal matching and retrieval stage, all the sentences in the case library will be used to calculate the word weights to construct word vectors by tfidf, and then the Lsi Model is used for dimensionality reduction to obtain dense vectors, so as to obtain the vectors corresponding to each document in the case library. At the same time, the input document is also transformed into a vector. Finally, the Sparse MatrixSimilarity class provided by Gensim is used to calculate the similarity between the two documents. Within the preset similarity threshold, the category corresponding to the document with the highest similarity to the text to be tested is used as the category of this text element, so as to obtain the corresponding key text information based on the case text with the best similarity. If it is lower than this similarity threshold, there is no relevant category, and it is prompted that the key text information cannot be recognized.
[0135] Furthermore, the method further includes step S5, which includes the following steps:
[0136] Establish corresponding association relationships based on each article title, the hierarchical relationship between different article titles, and the key text information corresponding to each article title.
[0137] Based on step S5, the following operations can be further performed according to the actual situation:
[0138] Based on each of the article titles, the hierarchical relationships between different article titles, and the key text sub-information corresponding to each article title, information restoration is performed to obtain a contract information extraction file corresponding to the contract document.
[0139] That is, the hierarchical relationships of different article titles in the contract document are obtained, and the text content corresponding to each article title is also obtained. Therefore, based on the article titles, combined with the hierarchical relationships and the corresponding text content, information restoration can be performed.
[0140] In a second aspect, as shown in Figure 3 the embodiments of the present application provide a contract information extraction device, which includes:
[0141] A file parsing module, which is used to receive a contract document, perform text parsing, and obtain text parsing data;
[0142] A layout restoration module, which is used to obtain multiple contract text paragraphs of the contract document based on the text parsing data, and is also used to obtain different levels of article titles of the contract document based on the text feature information corresponding to the contract text paragraphs;
[0143] An information extraction module, which is used to identify and obtain the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title; where
[0144] The text parsing data includes:
[0145] Text feature information, which is used to record the text content in the contract document;
[0146] Layout feature information, which is used to record the orientation corresponding to the text content;
[0147] Visual feature information, which is used to record the font and font size corresponding to the text content;
[0148] The text feature information includes multiple text feature sub-information, and the text feature sub-information is used to record any paragraph of text content in the contract document;
[0149] The text feature sub-information respectively corresponds to any level of the article title.
[0150] It should be noted that the text feature information corresponds to the literal content of the contract document, and the contract document can be specifically divided into literal contents of multiple paragraphs, denoted as contract text paragraphs. Therefore, when parsing, the literal content of the contract document can also be obtained paragraph by paragraph, denoted as different text feature sub-information. One text feature sub-information corresponds to one contract text paragraph, and each paragraph corresponds to article titles of different levels. Therefore, each of the text feature information corresponds to an article title, that is, each of the text feature information belongs to an article title corresponding to it;
[0151] In addition, due to the level characteristics of article titles of different levels, at least one contract text paragraph, that is, at least one text feature sub-information, is included under one article title;
[0152] When the content of the article title is too much, at least two or even multiple contract text paragraphs can be included under one article title, that is, two or even multiple text feature sub-information can be included.
[0153] It should be noted that the text feature information corresponds to the literal content of the contract document, and the contract document can be specifically divided into literal contents of multiple paragraphs. Therefore, when parsing, the literal content of the contract document can also be obtained paragraph by paragraph, denoted as different text feature sub-information, and each paragraph corresponds to article titles of different levels. Therefore, each of the text feature information corresponds to an article title, that is, each of the text feature information belongs to an article title corresponding to it.
[0154] Among them, when the file parsing module works specifically, a pdf parsing module can be trained using the contract document dataset, specifically an OCR recognition module trained based on the open-source project pdfminer combined with an OCR model;
[0155] For the input pdf file, the output is the contract parsing text, that is, the text parsing data, and its content at least includes text information, coordinates for indicating the position of the text, font, and font size.
[0156] In the embodiment of the present application, there is no need to write rules according to different layouts, extract contract information by identifying article titles and corresponding title levels, and combining corresponding key text information, which improves the generality of technical implementation while ensuring the accuracy of information extraction.
[0157] It should be noted that in the embodiment of the present application, the contract document can specifically be a file in pdf format.
[0158] Specifically, when the layout restoration module works, it can split the paragraphs in the contract document according to the text parsing data, so as to obtain multiple contract text paragraphs of the contract document;
[0159] The data basis for paragraph splitting can be the text feature information, layout feature information, and visual feature information of the text parsing data. When necessary, the corresponding punctuation marks can also be combined.
[0160] Specifically, when the layout restoration module works, based on the text feature information corresponding to the contract text paragraph, to obtain the article titles at different levels of the contract document, the following operations are included:
[0161] Based on the layout feature information and the visual feature information, obtain the article titles in the contract document from the text feature information;
[0162] Based on the layout feature information and the visual feature information, identify the hierarchical relationship between different article titles.
[0163] It should be noted that when the layout restoration module is specifically implemented, it includes at least two stages, and the specific situation is as follows:
[0164] The first stage is to identify the article titles in the document:
[0165] The training uses the content of each page of the pdf document as a sample unit, and adopts the longformer pre-trained model. The pre-trained base model is the publicly available Chinese model longformer-chinese-base-4096, which is optimized and trained on Roberta. By optimizing the self-attention structure of the transformer, the computational complexity is reduced, enabling the model to model long texts;
[0166] When necessary, the open-source pre-trained model longformer-chinese-base-4096 can also be retrained. The training corpus is the contract text in the industry, and the training method is MLM. Through further masked training, the model is made more sensitive to the vertical industry.
[0167] The input is text features, layout features, and visual features. The text features are mainly the text content in the pdf, which is input sentence by sentence. The layout features are the coordinate box information x0, y0, x1, y1 corresponding to the text box, which respectively correspond to the coordinates of the upper left corner and the lower right corner of the text box, and the height and width of each character are obtained through coordinate calculation. The visual features are information such as the font and font size of the text;
[0168] Finally, there are 9 features including token, x0, y0, x1, y1, width, height, fontname, and fontsize as inputs. The length of each feature is uniformly 2048, and padding is performed for those with insufficient length. By performing relevant calculations on the token, position and segment features can be generated. The three feature vectors of token, position, and segment are used as the input of the longformer, and a vector embedding1 with dimensions B*T*E will be obtained;
[0169] Round the coordinate data of x0, y0, x1, y1, width, and height. The index range is 0 - 1024. Vector embedding is obtained through random initialization to get embedding2; fontname and fontsize are indexed by constructing a dictionary using all the data, and vector embedding is obtained through random initialization to get embedding3. The three output vectors are weighted and summed, and then connected to a fully connected layer to construct a binary classifier to achieve article title recognition.
[0170] It should be noted that the contract file in pdf format includes multiple text boxes, which together constitute the text content of the pdf file.
[0171] The second stage: Identify the hierarchy of the titles in the document:
[0172] The longformer pre-trained model can also be used in this stage. The input data for this stage is the output data of the first stage. The difference is that the input of the first stage is all text, while the input of the second stage is the article title. The training uses the title of each pdf document as the sample unit, and the processing method of the features is the same, and the middle layer network structure is the same;
[0173] During the process of identifying the title and its hierarchy, relevant rules will be added to correct the model results. After the first stage of recognition is completed, the title sentences are grouped according to the corresponding font sizes, and then regular expressions are used to further search for the title lines with obvious identifiers in the text, such as sentences starting with "the first" or "1.1". After matching the relevant sentences, it will be checked whether they are in the set generated by the model. If they are, the sentence will be output as the title in the end; if not, it will not be output.
[0174] In the second stage, a similar method is also constructed. The output of the model is used to build a font size dictionary, where each level of the title corresponds to a unique font size. Regular expressions are used to search for all titles. The centered titles are usually the first-level titles, and the model results are partially corrected accordingly. Eventually, a hierarchical numbering like 1, 2, 2, 2, 3, 1, 1, 2 will be generated, which is then converted into a table of contents numbering structure like 1, 1.1, 1.2, 1.3, 1.3.1, 2, 3, 3.1. Additionally, title identifiers such as 1.19 and 2.3 usually correspond to the 19th and 3rd titles at that title level respectively, and the title numbers at that level can be corrected through this method.
[0175] Finally, the corrected results are stored in a key-value dictionary structure, where the key is the title and its corresponding level, and the value is the text content under the title.
[0176] It should be noted that in the second stage, the pipeline structure is used to identify the document titles in stages. This method avoids a large amount of rule writing, has strong portability, is not restricted by rules, and can greatly improve the accuracy and effect of the model when combined with manual rule correction.
[0177] Furthermore, the information extraction module is also used to obtain the corresponding text feature sub-information based on the contract text paragraph corresponding to the article title, and is also used to identify the text feature sub-information to identify the key text information corresponding to the article title; among them,
[0178] The text feature information includes multiple text feature sub-informations;
[0179] The text feature sub-information corresponds to a contract text paragraph.
[0180] Specifically, there is an implementation of the information extraction module, that is, based on the text feature information corresponding to the contract text paragraph corresponding to the article title, when identifying the key text information corresponding to the article title, it includes the following steps:
[0181] Based on the contract text paragraph corresponding to the article title, obtain the corresponding text feature sub-information corresponding to the contract text paragraph;
[0182] Identify the text feature sub-information to identify the key text information corresponding to the article title; among them,
[0183] The text feature information includes multiple text feature sub-informations;
[0184] The text feature sub-information corresponds to a contract text paragraph, which corresponds to the article title at any level.
[0185] Furthermore, the information extraction module is also used to obtain corresponding text feature sub-information of the contract text paragraph based on the contract text paragraph corresponding to the article title;
[0186] The information extraction module is also used to identify the text feature sub-information, compare it with the case sentences of different case texts in the preset case library, and compare the similarity with the case texts;
[0187] The information extraction module is also used to obtain the key text information corresponding to the article title based on the case text with the best similarity; where
[0188] The text feature information includes multiple text feature sub-informations;
[0189] The text feature sub-information corresponds to one contract text paragraph.
[0190] Specifically, there is another implementation of the information extraction module, that is, in the process of identifying and obtaining the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title, the following steps are included:
[0191] Obtain corresponding text feature sub-information of the contract text paragraph based on the contract text paragraph corresponding to the article title;
[0192] Identify the text feature sub-information, compare it with the case sentences of different case texts in the preset case library, and compare the similarity with the case texts;
[0193] Obtain the key text information corresponding to the article title based on the case text with the best similarity; where
[0194] The text feature information includes multiple text feature sub-informations;
[0195] The text feature sub-information corresponds to one contract text paragraph;
[0196] The text feature sub-information corresponds to any level of the article title.
[0197] It should be noted that in specific implementation, keyword fields will be preset, corresponding screening rules will be configured, and a method of fusing rules and algorithms will be adopted. If relevant rule methods such as regular expressions and artificial logic are not recognized, further matching and retrieval will be performed in the case library. The information extraction module includes the following two working conditions:
[0198] Case 1: The information corresponding to the keyword field will be directly set under the corresponding subtitle in some documents. For example, the confidentiality clause information will have a dedicated paragraph to elaborate on it, and its title is "Confidentiality Clause". For this category, the title can be directly indexed in the dictionary and output.
[0199] Case 2: The information corresponding to the keyword field is hidden in the text. For example, Party A: xxxxxx Company, Project Name: xxxxxx. For this situation, regular expressions are used for extraction and output.
[0200] For Case 2, in specific implementation, several sentences corresponding to this field will be input first, such as 6 - 10 sentences, and stored in the case library as cold start matching sentences. In the formal matching and retrieval stage, all the sentences in the case library will be used to calculate the word weights to construct word vectors by tfidf, and then the Lsi Model is used for dimensionality reduction to obtain dense vectors, so as to obtain the vectors corresponding to each document in the case library. At the same time, the input document is also transformed into a vector. Finally, the Sparse MatrixSimilarity class provided by Gensim is used to calculate the similarity between the two documents. Within the preset similarity threshold, the category corresponding to the document with the highest similarity to the text to be tested is used as the text element category of this text, so as to obtain the corresponding key text information based on the case text with the best similarity. If it is lower than this similarity threshold, there is no relevant category, and it will be prompted that the key text information cannot be recognized.
[0201] Furthermore, the layout restoration module is also used to obtain the article title in the contract document from the text feature information based on the layout feature information and the visual feature information.
[0202] The layout restoration module is also used to identify the hierarchical relationship between different article titles based on the layout feature information and the visual feature information.
[0203] Furthermore, the device also includes an association generation module, which is used to establish corresponding association relationships based on each article title, the hierarchical relationship between different article titles, and the key text information corresponding to each article title.
[0204] Based on the work content of the association generation module, the following operations can be further carried out according to the actual situation:
[0205] Based on each article title, the hierarchical relationship between different article titles, and the key text sub - information corresponding to each article title, information restoration is performed to obtain the contract information extraction file corresponding to the contract document.
[0206] That is, the hierarchical relationship of different article titles in the contract document is obtained, and the text content corresponding to each article title is also obtained. Therefore, based on the article titles, combined with the hierarchical relationship and the corresponding text content, information restoration can be carried out.
[0207] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0208] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for extracting contract information, characterized in that, The method includes the following steps: Receiving a contract document, performing text parsing, and obtaining text parsing data; Based on the text parsing data, obtaining multiple contract text paragraphs of the contract document; Based on the text feature information corresponding to the contract text paragraphs, obtaining different levels of article titles of the contract document, with at least one contract text paragraph under one article title; Based on the text feature information corresponding to the contract text paragraphs corresponding to the article title, identifying and obtaining the key text information corresponding to the article title; In the process of identifying and obtaining the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraphs corresponding to the article title, it includes a first process or a second process. The first process includes: Based on the contract text paragraphs corresponding to the article title, obtaining the corresponding text feature sub-information of the contract text paragraphs; Identifying the text feature sub-information and obtaining the key text information corresponding to the article title; The second process includes: Based on the contract text paragraphs corresponding to the article title, obtaining the corresponding text feature sub-information of the contract text paragraphs; Identifying the text feature sub-information, comparing it with the case sentences of different case texts in a preset case library, and comparing the similarity with the case texts; Based on the case text with the best similarity, obtaining the key text information corresponding to the article title; Wherein, the text parsing data includes: Text feature information, which is used to record the text content in the contract document; Layout feature information, which is used to record the orientation corresponding to the text content; Visual feature information, which is used to record the font and font size corresponding to the text content; The text feature information includes multiple text feature sub-informations; The text feature sub-information corresponds to one contract text paragraph; The text feature sub-information is used to record any paragraph of text content in the contract document; The text feature sub-informations respectively correspond to any level of article titles.
2. The contract information extraction method according to claim 1, wherein In the process of obtaining different levels of article titles of the contract document based on the text feature information corresponding to the contract text paragraphs, it includes the following steps: Based on the layout feature information and the visual feature information, obtaining the article titles in the contract document from the text feature information; Based on the layout feature information and the visual feature information, identifying the hierarchical relationship between different article titles.
3. The contract information extraction method according to claim 1, wherein, The method further includes the following steps: Based on each article title, the hierarchical relationship between different article titles, and the key text information corresponding to each article title, establishing a corresponding association relationship.
4. A contract information extraction device, characterized in that, The device includes: A file parsing module, which is used to receive a contract document, perform text parsing, and obtain text parsing data; A layout restoration module, which is used to obtain multiple contract text paragraphs of the contract document based on the text parsing data, and is also used to obtain different levels of article titles of the contract document based on the text feature information corresponding to the contract text paragraphs, with at least one contract text paragraph under one article title; An information extraction module, which is used to identify and obtain the key text information corresponding to the article title based on the text feature information corresponding to the contract text paragraph corresponding to the article title; The information extraction module is further used to obtain the text feature sub-information corresponding to the corresponding contract text paragraph based on the contract text paragraph corresponding to the article title, and is also used to identify the text feature sub-information to identify and obtain the key text information corresponding to the article title; The information extraction module is further used to obtain the text feature sub-information corresponding to the corresponding contract text paragraph based on the contract text paragraph corresponding to the article title The information extraction module is further used to identify the text feature sub-information and compare it with the case sentences of different case texts in a preset case library to compare the similarity with the case text; The information extraction module is further used to obtain the key text information corresponding to the article title based on the case text with the best similarity; wherein, The text parsing data includes: Text feature information, which is used to record the text content in the contract document; Layout feature information, which is used to record the orientation corresponding to the text content; Visual feature information, which is used to record the font and font size corresponding to the text content; The text feature information includes multiple text feature sub-informations, and the text feature sub-informations are used to record any paragraph of text content in the contract document; The text feature sub-informations respectively correspond to the article titles at any level; The text feature information includes multiple text feature sub-informations; The text feature sub-information corresponds to a contract text paragraph; The text feature sub-information is used to record any paragraph of text content in the contract document; The text feature sub-informations respectively correspond to the article titles at any level.
5. The contract information extraction device according to claim 4, wherein: The layout restoration module is further used to obtain the article title in the contract document from the text feature information based on the layout feature information and the visual feature information; The layout restoration module is further used to identify the hierarchical relationship between different article titles based on the layout feature information and the visual feature information.
Citation Information
Patent Citations
Information extraction method of nonstandard format documents of enterprises
CN106776538A
PDF full-automatic indexing system and method based on text features and grammatical rules
CN112307718A