A text element extraction method, device, equipment and medium
By generating candidate element labels and allowing users to select them when no matching model is found in the preset text element extraction model library, and combining this with a second text element extraction model, the problem of low efficiency in extracting new document layout elements is solved, achieving efficient and accurate text element extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies require a significant amount of manual annotation and model training time when dealing with new document formats, resulting in low efficiency of feature extraction algorithms.
By acquiring the text information and document format of the document to be identified, and using the unmatched model in the preset text feature extraction model library, the text information is input into the first text feature extraction model to generate candidate feature labels. The user selects the accurate label, and the labeled information is input into the second text feature extraction model to complete the feature extraction.
It achieves efficient optimization of the basic text element extraction model, improving the efficiency and accuracy of element extraction and reducing model training time.
Smart Images

Figure CN115690816B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for extracting text elements. Background Technology
[0002] In various stages of business operations, different types of contract documents circulate. Extracting key information from these documents requires time and manpower for identification, or pre-annotating large amounts of text and training an automatic element extraction model based on the annotation results. When new document formats need to be identified, significant time costs are incurred for manual annotation and model training. Summary of the Invention
[0003] This invention provides a method, apparatus, device, and medium for extracting text elements, which solves the problem of low efficiency in developing algorithmic models for extracting elements from new document formats. It achieves efficient optimization of basic text element extraction model algorithms, thereby improving the efficiency and accuracy of element extraction.
[0004] According to one aspect of the present invention, a method for extracting text elements is provided, the method comprising:
[0005] Obtain the file to be identified, and identify the text information and file format within the file;
[0006] When no matching text feature extraction model is found in the preset text feature extraction model library that corresponds to the document layout, the text information is input into the first text feature extraction model to obtain at least one candidate feature label for each text feature to be extracted.
[0007] Based on the user's label selection operation for the candidate element labels of each text element to be extracted, the accurate element label of each text element to be extracted is determined.
[0008] The text information with accurate feature labels is input into the second text feature extraction model to obtain the text feature model extraction results, thus completing the feature extraction process.
[0009] According to another aspect of the present invention, a text element extraction apparatus is provided, the apparatus comprising:
[0010] The text processing module is used to acquire the file to be recognized and to identify the text information and file format in the file.
[0011] The first text element extraction module is used to input text information into the first text element extraction model when no element extraction model corresponding to the document format is matched in the preset text element extraction model library, so as to obtain at least one candidate element label for each text element to be extracted.
[0012] The text label determination module is used to determine the accurate element label of each text element to be extracted based on the user's label selection operation on the candidate element labels of each text element to be extracted.
[0013] The second text feature extraction module is used to input text information with accurate feature labels into the second text feature extraction model, obtain the text feature model extraction results, and complete the feature extraction process.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory that is communicatively connected to at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the text element extraction method of any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the text element extraction method of any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring a file to be identified and recognizing its text information and file layout. When no matching text element extraction model corresponding to the file layout is found in a preset text element extraction model library, the text information is input into a first text element extraction model to obtain at least one candidate element label for each text element to be extracted. Based on the user's label selection operation for each candidate element label, the accurate element label for each text element to be extracted is determined. The text information labeled with the accurate element label is input into a second text element extraction model to obtain the text element model extraction result, thus completing the element extraction process. This technical solution solves the problem of low efficiency in developing algorithm models for element extraction from new file layouts, and achieves efficient optimization of basic text element extraction model algorithms, improving the efficiency and accuracy of element extraction.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a text element extraction method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart of another text element extraction method provided in an embodiment of the present invention;
[0024] Figure 3 This is a flowchart of another text element extraction method provided in the embodiments of the present invention;
[0025] Figure 4 This is a structural block diagram of a text element extraction device provided in an embodiment of the present invention;
[0026] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first" and "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Figure 1This is a flowchart illustrating a text element extraction method provided in an embodiment of the present invention. This embodiment is applicable to text element extraction scenarios, such as the extraction of text elements from various contracts. The method can be executed by a text element extraction device, which can be implemented in hardware and / or software, or configured in an electronic device.
[0030] like Figure 1 As shown, a text element extraction method includes the following steps:
[0031] S110. Obtain the file to be recognized and identify the text information and file format in the file.
[0032] Text information refers to the content of textual information in a document. Text information can be presented in the form of images, tables, and dialogues in the document to be recognized.
[0033] Document layout is used to determine the typesetting and categorization of text information within a document. For example, it can be a layout that includes text information presented in one or more formats, such as images, tables, and dialogues. Each document layout corresponds to a specific text element extraction model.
[0034] To facilitate computer processing, the text information in different presentation formats in the file to be recognized needs to be converted into content that the computer can recognize. After obtaining the file to be recognized, the text content in the file can be converted into corresponding text vectors by performing numerical conversions, converting non-numerical values to numerical values, filling in empty values, and other operations. Then, character recognition is performed to obtain the text information corresponding to the text content of the file to be recognized.
[0035] Null value imputation is based on statistical principles, filling in blanks according to the distribution of values of other objects in the initial dataset. Imputation methods include manual input, filling with specific values, filling with the average value, and the K-nearest neighbor method.
[0036] Special value filling treats null values as a special attribute value, unlike any other attribute value. For example, all null values can be filled with "unknown".
[0037] Mean imputation involves dividing the attributes in the initial dataset into numerical and non-numerical attributes and processing them separately. If a null value is numerical, it is filled with the average value of that attribute across all other objects. If a null value is non-numerical, it is filled with the mode principle from statistics. A similar method is conditional mean imputation. Conditional mean imputation does not take data from all objects in the dataset to calculate the mean; instead, it takes data from objects that share the same decision attribute value, filling in the missing attribute value with the most probable possible value.
[0038] K-nearest neighbor (K-means clustering) uses Euclidean distance or correlation analysis to determine the K nearest neighbors of a sample with missing data, and then takes a weighted average of the values of the K nearest neighbors to estimate the missing data of that sample.
[0039] S120. When no element extraction model corresponding to the document layout is matched in the preset text element extraction model library, the text information is input into the first text element extraction model to obtain at least one candidate element label for each text element to be extracted.
[0040] Understandably, it is necessary to set the corresponding text element extraction model according to the document layout to extract text elements. The preset text element extraction model library includes various document layouts and their corresponding text element extraction models.
[0041] Text feature extraction models are classification models used to categorize text information. They include various textual information and their corresponding feature labels. Feature labels represent the category of textual features, such as names, codes, locations, and quantities. Commonly used classification models include logistic regression, Naive Bayes, decision trees, Support Vector Machines (SVM), random forests, and gradient boosting trees. Logistic regression, also known as logistic regression analysis, has a computational cost that depends only on the number of features, resulting in fast training speeds and low memory usage, requiring only the storage of feature weights. For documents to be identified, their format and information can determine their subject matter, thus feature selection is needed to improve the classification efficiency of the Naive Bayes classifier. Decision trees are tree-like structures where each internal node represents a test on an attribute, each branch represents a test output, and each leaf node represents a category. They are inefficient for predicting continuous fields. Gradient boosting decision tree (GBDT) can be used to solve multi-class classification problems using the softmax function.
[0042] Specifically, when no matching element extraction model corresponding to the document layout is found in the preset text element extraction model library, the text information of the document to be identified cannot be directly classified based on the element extraction model in the text element extraction model library. Therefore, the text vector corresponding to the text information of the document to be identified needs to be input into the first text element extraction model to extract text elements and obtain one or more candidate element labels for each text element to be extracted for the user to select.
[0043] S130. Based on the user's label selection operation for the candidate element labels of each text element to be extracted, determine the accurate element label of each text element to be extracted.
[0044] Understandably, when the first text feature extraction model outputs one or more candidate features corresponding to the text information of the document to be identified, manual intervention is required to determine the feature label corresponding to the text information. Users can select feature labels through a visual human-computer interaction interface; if there is no label corresponding to the text feature among the candidate feature labels, users can also add a label as the feature label corresponding to the text information. The feature label selected or added by the user is used as the accurate feature label of the text feature to be extracted.
[0045] S140. Input the text information with accurate feature labels into the second text feature extraction model to obtain the text feature model extraction results and complete the feature extraction process.
[0046] The second text feature extraction model is used to extract features from text information based on accurate feature labels, resulting in text feature extraction results. For example, it can extract features from cells in a table within text information to obtain the features corresponding to the text information in that cell.
[0047] Understandably, the text information of the document to be identified, which is marked with accurate feature labels, is input into the second text feature extraction model. The second text feature extraction model extracts text features from the text information of the document to be identified corresponding to the label, and obtains the features corresponding to the text information in the document to be identified, thus completing the feature extraction process.
[0048] In an optional implementation, the method further includes: saving the second text feature extraction model that has completed the feature extraction process to a preset text feature extraction model library, and establishing a correspondence between the second text feature extraction model that has completed the feature extraction process and the document layout.
[0049] Understandably, it is necessary to save the second text element extraction model that has completed the element extraction process to the preset text element extraction model library, and establish a correspondence between the second text element extraction model that has completed the element extraction process and the document layout. The advantage of doing this is that when the text layout is recognized again, the relevant content can be directly read from the preset text element extraction model library to perform text element extraction, which reduces the training time of the text element extraction model and further improves the efficiency of text element extraction.
[0050] The technical solution of this invention involves acquiring a file to be identified and recognizing its text information and file layout. When no matching text element extraction model corresponding to the file layout is found in a preset text element extraction model library, the text information is input into a first text element extraction model to obtain at least one candidate element label for each text element to be extracted. Based on the user's label selection operation for each candidate element label, the accurate element label for each text element to be extracted is determined. The text information labeled with the accurate element label is input into a second text element extraction model to obtain the text element model extraction result, thus completing the element extraction process. This technical solution solves the problem of low efficiency in developing algorithm models for element extraction from new file layouts, and achieves efficient optimization of basic text element extraction model algorithms, improving the efficiency and accuracy of element extraction.
[0051] Figure 2 This is a flowchart of another text element extraction method provided by an embodiment of the present invention. This embodiment is applicable to text element extraction scenarios, particularly to text element extraction scenarios of various contracts. This embodiment belongs to the same inventive concept as the text element extraction method in the above embodiments, and further describes the process of obtaining the final extraction result before completing the element extraction process. This method can be executed by a text element extraction device, which can be implemented by software and / or hardware and integrated into an electronic device with application development capabilities.
[0052] like Figure 2 As shown, a text element extraction method includes the following steps:
[0053] S210. Obtain the file to be identified, identify the file format and file type in the file to be identified, and perform text format processing on the files to be identified for different file types to obtain the content of the target format file.
[0054] File types can be categories such as images, documents, and dialogues, used to describe files that share common properties and characteristics. Each file format includes content from one or more file types.
[0055] The target format file content is set for different file types. The preset format is obtained by processing the file to be recognized. For example, OCR (optical character recognition) is used to recognize the text in the image and / or the table in the document and convert it into text format. Alternatively, space filtering and character replacement are performed on the document and the dialogue to convert it into text format so that the computer can process it.
[0056] S220. Input the content of the target format file into the pre-trained BERT model to obtain the corresponding text vector matrix, which serves as the text information.
[0057] The target format file content is input into a pre-trained BERT model to extract features from the text content of the target format file, resulting in a corresponding text vector matrix, which serves as the text information.
[0058] The pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT) is a bidirectional encoding model. BERT uses the Transformer algorithm, which is based on the multi-head attention mechanism. BERT stacks multiple Transformer models and pre-trains bidirectional deep representations by adjusting the bidirectional Transformers in all layers. Moreover, the pre-trained BERT model can be configured with an additional output layer, making it widely applicable and reducing repetitive model training work.
[0059] The basic steps of the BERT model are as follows: First, randomly cover or replace any word or phrase in a sentence of text. Then, during pre-training, the model predicts which part is covered or replaced based on contextual understanding. In subsequent training, only the covered part is predicted. Next, select continuous contextual sentences from the report content and use the Transformer model to identify the continuity of sentences, thus realizing the BERT model for bidirectional prediction based on context. Finally, load the pre-trained BERT model and continue training on the target format file content.
[0060] The Transformer model enables parallel computation and has an encoder-decoder structure. The encoder module consists of multiple encoder layers, and the decoder module also consists of decoders with the same number of layers. Each encoder comprises two sub-layers: a Self-Attention layer and an FFN (Position-wise FeedForward Network). The encoder input first flows into the Self-Attention layer, which allows the encoder to use information from other words in the input sentence when encoding specific words. The output of the Self-Attention layer then flows into the feedforward network. The decoder also has a Self-Attention layer and an FFN, but there is an additional attention layer between them, which allows the decoder to focus on relevant parts of the target format file content.
[0061] S230. When no element extraction model corresponding to the document layout is matched in the preset text element extraction model library, the text information is input into the first text element extraction model to obtain at least one candidate element label for each text element to be extracted.
[0062] The first text element extraction model is a neural network model trained using Long Short-Term Memory (LSTM).
[0063] LSTM is a type of temporal recurrent neural network containing LSTM blocks or other neural network-like structures. It can memorize numerical values for indefinite periods. Each block has a gate that determines whether the input is memorized and whether it is output. Variations of LSTM include bidirectional recurrent neural networks (NRNNs) and deep recurrent neural networks (RNNs). The main structure of a bidirectional RNN consists of two unidirectional RNNs. At each time step, the input is simultaneously provided to these two RNNs in opposite directions, and the output is jointly determined by these two unidirectional RNNs. Deep RNNs, designed to enhance the model's expressive power, replicate the recurrent structure multiple times at each time step. The parameters are consistent across all layers, but can differ between layers.
[0064] S240. Based on the user's label selection operation for the candidate element labels of each text element to be extracted, determine the accurate element label of each text element to be extracted.
[0065] S250. Input the text information with accurate feature labels into the second text feature extraction model to obtain the text feature model extraction results and complete the feature extraction process.
[0066] The second text element extraction model can also be a neural network model trained based on long short-term memory.
[0067] The technical solution of this invention involves acquiring a file to be identified, identifying its format and type, and processing the text format for different file types to obtain target format file content. This target format file content is then input into a pre-trained BERT model to obtain a corresponding text vector matrix, which serves as the text information. When no matching text element extraction model for the file format is found in the preset text element extraction model library, the text information is input into a first text element extraction model to obtain at least one candidate element label for each text element to be extracted. Based on the user's label selection operation for each candidate element label, the accurate element label for each text element to be extracted is determined. The text information labeled with the accurate element label is then input into a second text element extraction model to obtain the text element model extraction result, thus completing the element extraction process. This technical solution, by processing the file content of the file to be identified to obtain the corresponding text vector matrix, solves the problem of low efficiency in developing algorithm models for element extraction of new file formats. It achieves efficient optimization of the basic text element extraction model algorithm, further improving the efficiency and accuracy of element extraction.
[0068] Figure 3 This is a flowchart of another text element extraction method provided by an embodiment of the present invention. This embodiment is applicable to scenarios involving text element extraction, especially for extracting text elements from various contracts. This embodiment belongs to the same inventive concept as the text element extraction method in the above embodiments, and further describes the process of extracting text elements according to pre-configured element extraction rules before completing the element extraction model. This method can be executed by a text element extraction device, which can be implemented by software and / or hardware and integrated into an electronic device with application development capabilities.
[0069] like Figure 3 As shown, a text element extraction method includes the following steps:
[0070] S310. Obtain the file to be recognized and identify the text information and file format in the file to be recognized.
[0071] S320. When no element extraction model corresponding to the document layout is matched in the preset text element extraction model library, the text information is input into the first text element extraction model to obtain at least one candidate element label for each text element to be extracted.
[0072] S330. Based on the user's label selection operation for the candidate element labels of each text element to be extracted, determine the accurate element label of each text element to be extracted.
[0073] S340. Input the text information marked with accurate feature labels into the second text feature extraction model to obtain the text feature model extraction results.
[0074] S350. Extract elements from the text information according to the pre-configured element extraction rules to obtain the text element rule extraction results.
[0075] In one optional implementation, the second text element extraction model is also pre-configured with element extraction rules for extracting elements from text information to obtain text element rule extraction results.
[0076] Furthermore, feature extraction is performed on the text information according to pre-configured feature extraction rules, including:
[0077] First, for the table in the document to be identified, the table's borders are identified and the position information of each cell in the table is determined. Based on the cell position information and the semantic information of the text within the cell, the elements are extracted.
[0078] Cell location information is used to indicate the cell's position in the file to be recognized; for example, it can be coordinates.
[0079] Specifically, firstly, the table in the document to be recognized is identified using an OCR (optical character recognition) algorithm to obtain the bounding box of the table and determine the position information of each cell in the table; then, element extraction is performed based on the cell position information and the semantic information of the text within the cell.
[0080] OCR algorithms refer to the process by which electronic devices (such as scanners or digital cameras) examine characters printed on paper and translate their shapes into computer text using character recognition methods. Specifically, it involves scanning text documents and then analyzing and processing image files to obtain text and layout information.
[0081] Using OCR algorithms for table recognition mainly involves three steps: text detection, text recognition, and table structure recognition.
[0082] (1) Text detection: Identify the coordinates of the text and the text box.
[0083] (2) Text recognition: Recognize the value in the text box, which is the text in the text box.
[0084] (3) Table structure recognition: Recognize the table structure and cell coordinates.
[0085] Specifically, feature extraction using OCR algorithms mainly includes the following steps:
[0086] 1. Image Input: Different image formats have corresponding storage formats and compression methods.
[0087] 2. Image preprocessing: mainly includes binarization, noise removal, character segmentation, and skew correction.
[0088] Binarization: The original image of the input orthopedic consumables is a color image. In order to speed up the recognition speed and improve the recognition efficiency, the content of the color image can be divided into foreground and background. The corresponding image only includes foreground information and background information. The foreground information can be defined as black and the background information as white to complete the binarization.
[0089] A common noise removal method is to use a one-dimensional discrete differential template to process the original image of orthopedic consumables in one direction or simultaneously in both the horizontal and vertical directions.
[0090] Character segmentation: This is used to solve the problem of character adhesion caused by the original image of orthopedic consumables. It requires an OCR analysis algorithm to segment the characters in the original image into images of individual characters.
[0091] Tilt correction: This is used to solve the problem of the original image being tilted. Tilt correction needs to be performed before character recognition to correct the orientation of the original image.
[0092] 3. Perform text detection: Perform text detection on the preprocessed text image, and identify the text box containing the text and the corresponding coordinate values of the text box through the text detection model.
[0093] Algorithms used for text detection mainly fall into two categories: regression-based and segmentation-based algorithms. Examples include regression-based algorithms such as EAST (Efficient and Accurate Scene Text) and TextBoxes, and segmentation-based algorithms such as DBNet, PAN, and FCENet.
[0094] 4. Text recognition: Character recognition can be performed through template matching to identify the text in the text boxes in the image and the coordinates of the corresponding text.
[0095] Specifically, character recognition is performed using template matching: First, a character template is created, which must have the same font format as the character to be recognized. Then, the character template is normalized, which improves processing speed. Finally, character recognition is performed by applying the same normalization process to the character to be recognized and the character template, iterating and comparing them. First, the difference between the character to be recognized and the character template is calculated, then the total pixel value of the resulting image is compared with a set character threshold. If the value is less than the set character threshold, it means that the character to be recognized and the template are the same character.
[0096] 5. Perform table structure recognition: identify the table borders and determine the position information of each cell in the table.
[0097] 6. Text Concatenation: First, aggregate the position information of each cell in the table with the coordinates of the text in the text boxes; second, sort the text boxes from top to bottom and from left to right; find the corresponding cell based on the coordinates of the text in the text box; then, concatenate the text within the cell to obtain the final text within the cell; finally, based on the pre-set correspondence between the position information of the table cells and the text within the cells, extract the elements from the cells containing text in the current table, completing the element extraction of the table in the file to be recognized.
[0098] Then, for the text content in the file to be identified, the text content is divided into paragraphs, and based on the paragraph division results, paragraph elements are extracted according to a preset regular expression matching strategy.
[0099] Regular expression matching strategies are implemented through regular expressions (regex), which are strings composed of characters and special symbols. They express a filtering logic for strings and divide paragraphs by matching a series of strings with similar characteristics according to a preset pattern. For example, a paragraph can be set to be a continuous line or multiple lines, with each line having at least one character and a newline character between lines. It can also be set to match whitespace characters.
[0100] Users can select regular expression matching strategies through a visual human-computer interaction interface. For example, they can select the splicing distance between preceding and following paragraphs, select the current paragraph of the text to splice with a set number of adjacent paragraphs, and the adjacent paragraphs include the paragraphs before and / or after the current paragraph. They can also choose to splice one or two paragraphs forward.
[0101] Specifically, for the text content in the file to be identified, users can select a regular expression matching strategy through a visual human-computer interaction interface, select the splicing distance between the preceding and following paragraphs, and match a series of strings with similar characteristics according to a preset pattern to achieve paragraph division and extract paragraph elements.
[0102] S360. When the extraction result of the text element model and the extraction result of the text element rule for any text element to be extracted are inconsistent, the final extraction result shall be determined from the extraction result of the text element model and the extraction result of the text element rule according to the preset result selection strategy.
[0103] The preset result selection strategy is to compare and deduplicate the results extracted by the text feature model and the results extracted by the text feature rules, remove the same content in the results extracted by the text feature model and the results extracted by the text feature rules, and find the inconsistent content in the results extracted by the text feature model and the results extracted by the text feature rules.
[0104] Specifically, first define a new array to store the unique elements extracted from the text feature model and the text feature rule extraction results; then iterate through each element extracted from the text feature model and compare it with each element extracted from the text feature rule; if there are no duplicate elements, store the element in the new set; if the length of the new set is equal to the loop variable after the loop ends, it means that there are no duplicate elements in the new set.
[0105] It is understandable that the text feature extraction results obtained by the text feature extraction model for the text information of the document to be identified may differ from the text feature extraction results obtained according to the pre-configured feature extraction rules. When the results are inconsistent, the text feature extraction results obtained according to the pre-configured feature extraction rules will be used as the final extraction results.
[0106] Furthermore, the method also includes: performing format and / or content correction processing on the final element extraction results of the document to be identified.
[0107] Understandably, it is necessary to verify and modify the final element extraction results of the document to be identified. For example, verification and modification can be performed on the spelling of English words, the capitalization of English words, punctuation marks, personal names, place names, and the capitalization of Chinese numbers.
[0108] The technical solution of this invention involves acquiring a file to be identified and recognizing its text information and file format. When no matching text element extraction model corresponding to the file format is found in a preset text element extraction model library, the text information is input into a first text element extraction model to obtain at least one candidate element label for each text element to be extracted. Based on the user's label selection operation for each candidate element label, the accurate element label for each text element to be extracted is determined. The text information labeled with the accurate element label is input into a second text element extraction model. Element extraction is performed on the text information according to pre-configured element extraction rules to obtain text element rule extraction results. When the text element model extraction result and the text element rule extraction result for any text element to be extracted are inconsistent, the final extraction result is determined from the text element model extraction result and the text element rule extraction result according to a preset result selection strategy. This technical solution solves the problem of low efficiency in developing algorithm models for element extraction of new file formats, and achieves efficient optimization of basic text element extraction model algorithms, further improving the efficiency and accuracy of element extraction.
[0109] Figure 4 This is a structural block diagram of a text element extraction device provided in an embodiment of the present invention. This embodiment is applicable to text element extraction scenarios. The device can be implemented in hardware and / or software and integrated into an electronic device with application development capabilities.
[0110] like Figure 4 As shown, a text element extraction device includes: a text processing module 410, a first text element extraction module 420, a text label determination module 430, and a second text element extraction module 440.
[0111] The text processing module 410 is used to acquire the file to be recognized and to recognize the text information and file format in the file to be recognized.
[0112] The first text element extraction module 420 is used to input text information into the first text element extraction model when no element extraction model corresponding to the document format is matched in the preset text element extraction model library, so as to obtain at least one candidate element label for each text element to be extracted.
[0113] The text label determination module 430 is used to determine the accurate element label of each text element to be extracted based on the user's label selection operation on the candidate element labels of each text element to be extracted.
[0114] The second text feature extraction module 440 is used to input text information marked with accurate feature labels into the second text feature extraction model, obtain the text feature model extraction results, and complete the feature extraction process.
[0115] The technical solution of this invention involves acquiring a file to be identified and recognizing the text information and file format within it. When no matching element extraction model corresponding to the file format is found in a preset text element extraction model library, the text information is input into a first text element extraction model to obtain at least one candidate element label for each text element to be extracted. Based on the user's label selection operation for each candidate element label, the accurate element label for each text element to be extracted is determined. The text information marked with the accurate element label is input into a second text element extraction model to obtain the text element model extraction result, thus completing the element extraction process. This solution solves the problem of low efficiency in developing algorithm models for element extraction of new file formats, and achieves efficient optimization of basic text element extraction model algorithms, improving the efficiency and accuracy of element extraction.
[0116] Optionally, the device is used for:
[0117] Save the second text feature extraction model that has completed the feature extraction process to the preset text feature extraction model library, and establish the correspondence between the second text feature extraction model that has completed the feature extraction process and the document layout.
[0118] Optionally, the text processing module 410 is used for:
[0119] Identify the file type in the file to be identified, and perform text format processing on the file to be identified for different file types to obtain the content of the target format file;
[0120] The content of the target format file is input into the pre-trained BERT model to obtain the corresponding text vector matrix, which serves as the text information.
[0121] Optionally, the second text feature extraction module 440 is used for:
[0122] The text information is extracted according to the pre-configured feature extraction rules to obtain the text feature rule extraction results;
[0123] When the extraction results of any text element model and the extraction results of text element rules are inconsistent, the final extraction result is determined from the extraction results of text element model and text element rules according to the preset result selection strategy.
[0124] Optionally, the second text feature extraction module 440 is also used for:
[0125] For tables in the document to be recognized, the table outlines are identified and the position information of each cell in the table is determined. Based on the cell position information and the semantic information of the text within the cell, the elements are extracted.
[0126] For the text content in the file to be identified, the text content is divided into paragraphs, and based on the paragraph division results, paragraph elements are extracted according to a preset regular expression matching strategy.
[0127] Optionally, the device is also used for:
[0128] The final element extraction results of the document to be identified are processed for format and / or content correction.
[0129] The text element extraction device provided in this embodiment of the invention can execute a text element extraction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0130] Figure 5 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, or other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), or other similar computing devices. The components shown herein, their connections, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0131] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0132] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as a keyboard or mouse; output unit 17, such as various types of displays or speakers; storage unit 18, such as a disk or optical disk; and communication unit 19, such as a network card, modem, or wireless transceiver. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0133] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), any suitable processor, controller, or microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a text feature extraction method.
[0134] In some embodiments, a text feature extraction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the text feature extraction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a text feature extraction method by any other suitable means (e.g., by means of firmware).
[0135] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0136] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on the machine or partially on the machine, or as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0137] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0139] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0140] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0141] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0142] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A text element extraction method characterized by, The method comprises the following steps: obtaining a to-be-recognized file, and recognizing text information and a file format in the to-be-recognized file; when no element extraction model corresponding to the file format is matched in a preset text element extraction model library, inputting the text information into a first text element extraction model to obtain at least one candidate element label of each to-be-extracted text element; determining accurate element labels of the to-be-extracted text elements according to label selection operations of a user on the candidate element labels of the to-be-extracted text elements; inputting the text information marked with the accurate element labels into a second text element extraction model to obtain a text element model extraction result, and completing an element extraction process; before the element extraction process is completed, the method further comprises the following steps: performing element extraction on the text information according to a preconfigured element extraction rule to obtain a text element rule extraction result; when the text element model extraction result and the text element rule extraction result of any to-be-extracted text element are inconsistent, determining a final extraction result in the text element model extraction result and the text element rule extraction result according to a preset result selection strategy; wherein the preset result selection strategy is to compare and remove the same content in the text element model extraction result and the text element rule extraction result, and find the inconsistent content in the text element model extraction result and the text element rule extraction result.
2. The method of claim 1, wherein, The method further comprises the following steps: saving the second text element extraction model completing the element extraction process to the preset text element extraction model library, and establishing a corresponding relationship between the second text element extraction model completing the element extraction process and the file format.
3. The method of claim 1, wherein, The method further comprises the following steps: recognizing a file type in the to-be-recognized file, and performing text format processing on the to-be-recognized file of different file types to obtain target format file content; inputting the target format file content into a pre-trained BERT model to obtain a corresponding text vector matrix as the text information.
4. The method of claim 1, wherein, The first text element extraction model and the second text element extraction model are neural network models trained based on a long short-term memory recurrent neural network.
5. The method of claim 4, wherein, The method further comprises the following steps: for a table in the to-be-recognized file, recognizing a frame line of the table and determining position information of each cell in the table, and performing element extraction based on the position information of the cell and text semantic information in the cell; for text content in the to-be-recognized file, performing paragraph division on the text content, and performing paragraph element extraction according to a preset regular matching strategy based on a paragraph division result.
6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises the following steps: performing format and / or content correction processing on the final element extraction result of the to-be-recognized file.
7. A text element extraction apparatus characterized by comprising: The method comprises the following steps: a text processing module is configured to obtain a to-be-recognized file, and recognize text information and a file format in the to-be-recognized file; The first text element extraction module is configured to input the text information into a first text element extraction model to obtain at least one candidate element label of each text element to be extracted when no element extraction model corresponding to the file format is matched in the preset text element extraction model library. The text label determination module is configured to determine accurate element labels of the text elements to be extracted according to a label selection operation of a user on the candidate element labels of the text elements to be extracted. The second text element extraction module is configured to input the text information marked with the accurate element labels into a second text element extraction model to obtain a text element model extraction result and complete an element extraction process. The second text element extraction module is configured to: extract elements from the text information according to a preconfigured element extraction rule to obtain a text element rule extraction result; when the text element model extraction result and the text element rule extraction result of any text element to be extracted are inconsistent, determine a final extraction result in the text element model extraction result and the text element rule extraction result according to a preset result selection strategy, wherein the preset result selection strategy is to compare and remove the same content in the text element model extraction result and the text element rule extraction result, and find the inconsistent content in the text element model extraction result and the text element rule extraction result.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the text element extraction method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the text element extraction method of any one of claims 1-6 when executed by the processor. The computer readable storage medium stores computer instructions for enabling the processor to execute the text element extraction method of any one of claims 1-6 when executed by the processor.
Citation Information
Patent Citations
Text recognition method, device and equipment and storage medium
CN111967437A
Rapid and automatic table extraction method based on deep neural network
CN112883795A
Document information intelligent acquisition and error correction method and system based on OCR technology and equipment
CN113936130A
Text marking method and device, electronic equipment and computer readable storage medium
CN114548051A