Information extraction method and device for bill image, storage medium and electronic equipment
By identifying bill images and matching them with tree-structured template configuration files, bill information can be accurately extracted, solving the problem of low efficiency in paper bill information extraction in the existing technology and achieving more efficient information extraction.
Patent Information
- Application Number
- CN202210641781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-06-08
AI Technical Summary
The efficiency of paper bill information extraction in the existing technology is low, especially due to the poor generalization of the model and label mismatch problems.
By identifying the target bill image, a set of recognized text boxes is obtained. Based on the set of multi-predicate template configuration files associated with the tree structure, template configuration files with high matching degrees are screened out to accurately extract bill information.
The efficiency and accuracy of paper bill information extraction are improved, and the problem of low information extraction efficiency in the prior art is solved.
Smart Images

Figure CN115063784B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, in particular to a bill image information extraction method and device, a storage medium and an electronic device. BACKGROUND
[0002] Structured entry of paper bills is a prerequisite for many further analysis based on bills, and related technologies extract paper bill information through machine learning based algorithms. In particular, in existing deep learning, structured information extraction is treated as a sequence named entity recognition (NER) task in natural language processing (NLP) or a method of classifying nodes using a graph neural network (GNN). The above method requires a relatively high accuracy of the model, and requires a large amount of data labeling, and the label categories of different bills may not be consistent, such as name, gender, and ID number labels for an ID card, and no train number label for a train ticket. The training set is prone to label mismatch problems when applied, and the model has poor generalization, so the efficiency of extracting information from paper bills is low. SUMMARY
[0003] The embodiments of the present application provide a bill image information extraction method and device, a storage medium and an electronic device to at least solve the technical problem of low efficiency of paper bill information extraction.
[0004] According to an aspect of an embodiment of the present application, a bill image information extraction method is provided, comprising: identifying a target bill image to obtain a set of identified text boxes and text information in each identified text box in the set of identified text boxes; based on the text information in each identified text box, obtaining a first reference template configuration file from a set of template configuration files; wherein each template configuration file in the set of template configuration files is a template file obtained by annotating the type and field of a bill image, and the set of template configuration files includes multiple predicates associated by a tree structure; filtering the first reference template configuration file to obtain a second reference template configuration file; wherein the text information of the positioning field in the template configuration file in the second reference template configuration file contains the text information in the identified text box; and determining a target template configuration file according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box.
[0005] According to another aspect of the embodiments of the present application, there is also provided an information extraction device for a bill image, comprising: an identification unit configured to identify a target bill image to obtain a set of identified text boxes and text information in each identified text box in the set of identified text boxes; an acquisition unit configured to acquire a first reference template configuration file from a set of template configuration files based on the text information in each identified text box; wherein each template configuration file in the set of template configuration files is a template file obtained by field annotation according to a bill image, and the set of template configuration files comprises multiple predicate associated by a tree structure; a screening unit configured to screen the first reference template configuration file to obtain a second reference template configuration file; wherein the text information of a positioning field in the template configuration file in the second reference template configuration file contains the text information in the identified text box; a determination unit configured to determine a target template configuration file according to a matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box; and an extraction unit configured to extract bill information in the target bill image according to the target template configuration file.
[0006] According to still another aspect of the embodiments of the present application, there is also provided an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the information extraction method for a bill image by using the computer program.
[0007] According to still another aspect of the embodiments of the present application, there is also provided a computer readable storage medium storing a computer program, wherein the computer program is configured to execute the information extraction method for a bill image when running.
[0008] In the embodiment of the present application, the target bill image is recognized to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes; a first reference template configuration file is obtained from a set of template configuration files based on the text information in each recognized text box; each template configuration file in the set of template configuration files is a template file obtained by annotating types and fields according to bill images, and the set of template configuration files includes multiple predicates associated by a tree structure; the first reference template configuration file is filtered to obtain a second reference template configuration file; the text information of a positioning field in the template configuration file in the second reference template configuration file contains the text information in the recognized text box; a target template configuration file is determined according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each recognized text box; and the target bill image is extracted according to the target template configuration file. In the above method, since a target bill template with a high matching degree is filtered from the bill template with annotated types and regions, and then the information in the bill image is extracted according to the target bill template, not only the information in the bill image can be accurately obtained, but also the efficiency of extracting information of a paper bill can be improved, thereby solving the technical problem of low efficiency of extracting information of a paper bill. BRIEF DESCRIPTION OF DRAWINGS
[0009] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0010] Figure 1 FIG. 1 is a schematic diagram of an application environment of an optional bill image information extraction method according to an embodiment of the present application;
[0011] Figure 2 FIG. 2 is a schematic diagram of an application environment of another optional bill image information extraction method according to an embodiment of the present application;
[0012] Figure 3 FIG. 3 is a flowchart of a bill image in a related art according to an embodiment of the present application;
[0013] Figure 4 FIG. 4 is a schematic diagram of bill image annotation according to an embodiment of the present application;
[0014] Figure 5 FIG. 5 is a flowchart of another optional bill image information extraction method according to an embodiment of the present application;
[0015] Figure 6is a flowchart of another optional information extraction method of a bill image according to an embodiment of the present application;
[0016] Figure 7 is a flowchart of another optional information extraction method of a bill image according to an embodiment of the present application;
[0017] Figure 8 is a flowchart of another optional information extraction method of a bill image according to an embodiment of the present application;
[0018] Figure 9 is a flowchart of another optional information extraction method of a bill image according to an embodiment of the present application;
[0019] Figure 10 is a structural diagram of an optional information extraction device of a bill image according to an embodiment of the present application;
[0020] Figure 11 is a structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by the personnel in the art without creative labor should belong to the protection scope of the present application.
[0022] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] According to an aspect of an embodiment of the present application, there is provided a bill image information extraction method. Optionally, as an optional implementation, the above-mentioned bill image information extraction method can be applied to, but is not limited to, a bill image information extraction method as shown in the following table: Figure 1The application environment shown in FIG. This application environment includes a terminal device 102 for human-computer interaction with a user, a network 104, and a server 106. User 108 can interact with terminal device 102, which runs a bill image information extraction application. Terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. The human-computer interaction screen 1022 is used to display multiple bill images; the processor 1024 is used to acquire a target bill image. The memory 1026 is used to store the multiple bill images.
[0024] In addition, the server 106 includes a database 1062 and a processing engine 1064. The database 1062 is used to store the multiple bill images. The processing engine 1064 is configured to recognize a target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set; obtain a first reference template configuration file from a set of template configuration files based on the text information in each recognized text box; wherein each template configuration file in the set of template configuration files is a template file obtained by annotating the bill image with a type and a field, and the set of template configuration files includes multiple predicates associated via a tree structure; filter the first reference template configuration file to obtain a second reference template configuration file; wherein the text information in the positioning field in the template configuration file in the second reference template configuration file includes the text information in the recognized text box; determine a target template configuration file based on the degree of match between each reference configuration file in the second reference configuration file and the text information in each recognized text box; extract the bill information from the target bill image according to the target template configuration file; and send the extracted bill information to the client of the terminal device 102.
[0025] In one or more embodiments, the information extraction method of the bill image described above can be applied to Figure 2 In the application environment shown. Figure 2 As shown, human-computer interaction can be performed between user 202 and user device 204. User device 204 includes memory 206 and processor 208. In this embodiment, user device 204 can, but is not limited to, refer to and perform the operations performed by the terminal device 102 to obtain bill information in the bill image.
[0026] Optionally, the terminal device 102 and the user equipment 204 include, but are not limited to, a mobile phone, a tablet computer, a notebook computer, a PC, a vehicle-mounted electronic device, a wearable device, and the like. The network 104 can include, but is not limited to, a wireless network or a wired network. The wireless network includes WIFI and other wireless communication networks. The wired network can include, but is not limited to, a wide area network, a metropolitan area network, and a local area network. The server 106 can include, but is not limited to, any hardware device capable of computing. The server can be a single server, a server cluster composed of multiple servers, or a cloud server. The above is only an example, and the present embodiment does not make any limitation on this.
[0027] The structured input application scenario of paper bills is relatively wide, and is also a prerequisite for further analysis based on bills. In related technologies, the following structured information extraction algorithms mainly exist:
[0028] Firstly, based on human-defined rules, this method excessively depends on the good or bad of rule customization, and different rules need to be defined for each bill, which has low universality and is difficult to maintain.
[0029] Secondly, the algorithm based on machine learning. In particular, in existing deep learning, the structured information extraction is treated as a sequence named entity recognition (NER) task in natural language processing (NLP) or a method of classifying nodes using a graph neural network (GNN). This method requires a high accuracy of the model and a large amount of data labeling. Moreover, the label categories of different bills may not be consistent, such as the identity card having labels such as name, gender, and ID number, while the train ticket does not have labels such as train number, which causes the training set and the application to have a label mismatch problem, and the generalization of the model is not strong.
[0030] In addition, the template matching method. The existing template matching has many limitations. The positioning field can only mark the non-repeated field on the text, the extracted field category is limited, and the extracted content is single. Only fixed fields can be extracted. When there are many template files, the real-time performance and accuracy of matching are not high, which further limits the application scenario of the template matching method.
[0031] In order to solve the above technical problems, as an optional implementation manner, as shown in Figure 3 The information extraction method of the bill image provided by the embodiment of the present application includes the following steps:
[0032] S302, performing recognition on the target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes.
[0033] In the embodiment of the present application, first, different paper bills can be photographed or scanned to extract corresponding bill images, and then the bill images are recognized. Here, the text box detection of the target bill image includes but is not limited to the detection-based method and the segmentation-based method, for example, using the segmentation-based deep learning method DB-Net to perform text segmentation, and obtaining the BBox (Bounding Box) where the text is located, i.e. the text box, through post-processing, which includes text information. When performing text recognition on the above-mentioned text box, it includes but is not limited to the algorithm based on CTC-Loss loss function and attention mechanism, such as using CRNN deep learning network to perform text recognition on the detected BBox to obtain the text information in the text box.
[0034] S304, based on the text information in each recognized text box, obtaining a first reference template configuration file from a set of template configuration files; wherein each template configuration file in the set of template configuration files is a template file obtained by annotating the type and field of the bill image, and the set of template configuration files includes multiple predicates associated through a tree structure.
[0035] Specifically, as shown in Figure 4 Each template picture in the template configuration file includes an annotated positioning field (the area where the five-pointed star is located), an annotated fixed field (the area where the circle is located), an annotated linearly associated area (the rectangular area), an annotated text block area (such as medical advice), and an annotated image area (medical image), etc. The set of template configuration files includes a tree bifurcation according to a certain specific word to narrow the scope of the template or improve the discrimination of similar templates. For example, taking the first hospital as the specific word, i.e. the root node of the tree structure, then taking this root node as the basis, then taking the text in all or part of the text box as a text string T, such as T = {first hospital, test report…audit time}, etc., then taking the test report as a screening condition, taking the set of template configuration files containing the test report in the set of template configuration files as the left child node, and taking the set of template configuration files not containing the test report as the right child node, and finally taking the template configuration file set containing the audit time or not as the root node. In this way, the template configuration files can be classified and aggregated according to different text words. For example, the first reference template configuration file is a template set {t1, t2,... tk} obtained according to T = {first hospital, test report…audit time}.
[0036] S306, screening the first reference template configuration file to obtain a second reference template configuration file; wherein text information of a positioning field in a template configuration file in the second reference template configuration file contains text information in an identified text box.
[0037] In the embodiment of the application, for example, the text center coordinates and text content of the positioning field in the template configuration file are [(x1(1), y1(1), t1(1)), (x1(2), y1(2), t1(2),..., (x1(m), y1(m), t1(m))], and the text center coordinates and text content of the detected text in the target bill image are [(x2(1), y2(1), t2(1)), (x2(2), y2(2), t2(2),..., (x2(n), y2(n), t2(n))]. The Cartesian product of the template text and the detected text is traversed, if the text content of t2(i) contains the text content of t1(j), it is considered that the ith detected text in the target bill image and the jth template text are possibly matched, that is, the second reference template configuration file.
[0038] S308, determining a target template configuration file according to a matching degree of each reference configuration file in the second reference configuration file and text information in each identified text box.
[0039] The matching number (n) and the matching error (err) of the to-be-identified bill and each template are calculated, the matching error is defined as a distance error of a matching point after a perspective transformation of a template file matching pair to the to-be-identified bill, the distance error can be selected as a Euclidean distance, and the greater the distance error, the greater the matching error. The number of positioning points of the template is denoted as N, and the matching rate ratio is defined as n / N. The optimal template is selected in the following manner:
[0040] Suppose that the Choice field in the above step is configured with a selection function fun(ratio, err), the optimal template is selected according to the configured selection manner, that is, the matched template; otherwise, a default manner is used: the templates with ratio>r0 (r0 is a defined constant) are filtered through a threshold, and then the template with the minimum matching error err is selected as the matched template from the templates meeting the threshold filtering condition.
[0041] S310, extracting bill information in the target bill image according to the target template configuration file.
[0042] Extract fixed field. Project the labeled fixed field using the calculated perspective transformation matrix, find the nearest text box around, the nearest text box as the extraction box, and the text content as the extracted content. If the distance is greater than the threshold or the text box is not intersected, it is considered to be missed, and the text content can be re-recognized at the text box position.
[0043] Extract linearly associated region information. Project the labeled linearly associated region using the above calculated perspective transformation matrix to obtain the linear region. Further filter the text box to extract the text box intersecting the linear region. Connect the text lines of the text box, specifically, select the text box from left to right and top to bottom, and construct a straight line with the center y coordinate of the text box to the right. The text box passed by the straight line is considered as the same text line. In this way, all text lines are connected. According to the attribute of each region labeled, the attribute of each text content of the text line is assigned to obtain a list of all (key, val, unit) matching tuples.
[0044] Extract text block region information. Project the labeled text block region using the above calculated perspective transformation matrix, and splice the text block content intersecting the text box.
[0045] Extract image region information. Project the labeled text block region using the above calculated perspective transformation matrix, and further process the projected region according to the labeled process, such as image classification, segmentation, etc.
[0046] In the embodiment of the present application, the target bill image is recognized to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes; a first reference template configuration file is obtained from a set of template configuration files based on the text information in each recognized text box; each template configuration file in the set of template configuration files is a template file obtained by annotating types and fields according to a bill image, and the set of template configuration files includes multiple predicates associated by a tree structure; the first reference template configuration file is filtered to obtain a second reference template configuration file; the text information of the positioning field in the template configuration file in the second reference template configuration file contains the text information in the recognized text box; a target template configuration file is determined according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each recognized text box; and the target bill image is extracted according to the target template configuration file. In the above method, since a target bill template with a high matching degree is filtered from the bill template with annotated types and regions, and then the information in the bill image is extracted according to the target bill template, not only the information in the bill image can be accurately obtained, but also the efficiency of extracting information of a paper bill can be improved, thereby solving the technical problem of low efficiency of extracting information of a paper bill.
[0047] In one or more embodiments, the recognizing the target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes comprises:
[0048] The target bill image is processed based on an image segmentation model to obtain the set of recognized text boxes.
[0049] Specifically, for example, text segmentation is performed using a deep learning model based on segmentation, DB-Net, for example, but not limited to, and a BBox (Bounding Box) in which the text is located, i.e., a recognized text box, is obtained through post-processing.
[0050] The set of text boxes is recognized by a character recognition model to obtain the text information in each recognized text box.
[0051] Specifically, text recognition is performed on the detected recognized text box. For example, but not limited to, algorithms based on CTC-Loss loss function and attention mechanism, such as a CRNN deep learning network, are used to perform text recognition on the detected BBox.
[0052] In one or more embodiments, before the target bill image is recognized, the method further comprises:
[0053] The regions where various types of information in the bill template picture are located are labeled by types;
[0054] The labeled bill template picture is divided into the template configuration file set including the multi-predicate according to the tree structure based on the preset word as the root node and the text content in the bill template picture.
[0055] As shown in Figure 4 The preset word is the first hospital, and the template configuration file containing other recognized text in the bill template image is selected from the template configuration file set based on the word as the root node. The multi-predicate contains all or part of the recognized text string in the bill template image.
[0056] In one or more embodiments, the labeling of the regions where various types of information in the bill template picture are located includes at least one of the following:
[0057] The positioning field in the bill template picture is labeled to obtain a set of labeled positioning fields; wherein the positioning field is a field with fixed relative position and text content in the bill template picture;
[0058] The fixed field in the bill template picture is labeled to obtain a set of labeled fixed fields; wherein the fixed field is a field with fixed relative position in the bill template picture;
[0059] The linearly associated tuple in the bill template picture is labeled to obtain a set of labeled linearly associated tuples; wherein the linearly associated tuple includes a plurality of associated fields;
[0060] The text block region in the bill template picture is labeled to obtain a set of labeled text block regions;
[0061] The image region in the bill template picture is labeled to obtain a set of labeled images.
[0062] In the embodiments of the present application, it mainly includes labeling positioning fields, labeling fixed fields, labeling linearly associated regions, labeling text block regions, etc. In addition to the positioning field, the others are optional labels, as shown in Figure 4 The region where the five-pointed star is located is the positioning field, the region where the circle is located is the fixed field, and the region where the rectangle is located is the linearly associated region. The labeling types are as follows:
[0063] Positioning field: the positioning field is to extract the information of the picture and match it with the existing template. Since the perspective transformation of the picture is involved, the positioning field must select the field with fixed relative position and text content in the picture, and at least four fields are included. The selected fields are distributed in the largest possible area of the picture, such as the four corners of the picture. The labeling method is as follows: the labeling type is selected as the positioning field, and the text content is the corresponding text content, to obtain the positioning field set {Xn} (n>=4).
[0064] Fixed field: the fixed field is a field with fixed relative position to be recognized, and the text content is the field to be extracted. For example, the name and gender fields on the medical test sheet. The labeling method is as follows: the labeling type is selected as the fixed field, and the labeling attribute key is the corresponding key to be recognized.
[0065] Linearly associated region: the linearly associated region is a plurality of associated regions, such as the test item column, test result column, and unit column on the medical test sheet. The labeling method is as follows: the labeling type is selected as the linearly associated region, and the labeling attribute corresponds. For example, the item is labeled as key-1, the result is labeled as val-1, and the unit is labeled as unit-1. If there are two columns, the second column is labeled as key-2, val-2, and unit-2, and so on.
[0066] Text block region: the text block region is a whole block of text content to be recognized, such as the medical order information on the medical bill. The labeling method is as follows: the labeling type is selected as the text block region, and the attribute key is the key to be recognized.
[0067] Picture region: the picture region is the picture information to be processed, such as some medical images. The labeling method is as follows: the labeling type is selected as the picture region, and the attribute is the processing method required by the picture and the allocated structured key.
[0068] In one or more embodiments, the first reference template configuration file is obtained from the template configuration file set based on the text information in each recognized text box, comprising:
[0069] Determine the initial matching field according to the text information in all recognized text boxes;
[0070] Take the initial matching field as the root node, and take the remaining fields of all recognized text boxes as the child nodes in turn, to traverse each multi-element predicate in the template configuration file set, to obtain the first reference template configuration file.
[0071] Specifically, based on the initial matching field for identifying any word in the text box, the initial matching field is used to enter the root node, and the selection of the sub-tree is performed according to different predicates, for example, the set of words containing the current predicate is taken as the left sub-tree (left child node set), and the set of words not containing the current predicate is taken as the right sub-tree (left child node set), and thus, according to the traversal result, a template set of multiple multi-predicate groups can be obtained, the template set is taken as a candidate template set, and the selection manner of the candidate template set and the optimal template is output.
[0072] In one or more embodiments, the screening of the first reference template configuration file to obtain the second reference template configuration file comprises:
[0073] Obtaining a set of positioning fields contained in all template configuration files in the first reference template configuration file;
[0074] Determining a Cartesian product of the text information in each identified text box and the set of positioning fields in the first reference template configuration file;
[0075] According to the reference template configuration file corresponding to the same ordered pair text in the Cartesian product, the second reference template configuration file is screened from the first reference template configuration file.
[0076] In one or more embodiments, the screening of the second reference template configuration file from the reference template configuration file corresponding to the same ordered pair text in the Cartesian product comprises:
[0077] Determining an offset vector corresponding to the text boxes of the two texts of the ordered pair of the same text;
[0078] According to the directionality and the modulus length of the offset vector, the filtered reference template configuration file is obtained through a box plot.
[0079] For example, the text center coordinates and the text content of the positioning fields in the template configuration file are: [(x1(1), y1(1), t1(1)), (x1(2), y1(2), t1(2),..., (x1(m), y1(m), t1(m))], and the text center coordinates and the text content of the corresponding detected text in the target bill image are: [(x2(1), y2(1), t2(1)), (x2(2), y2(2), t2(2),..., (x2(n), y2(n), t2(n))]. The Cartesian product of the template text and the detected text is traversed, and if the text content of t2(i) contains the text content of t1(j), it is considered that the i-th detected text in the target bill image and the j-th template text are possibly matched, that is, the second reference template configuration file.
[0080] The offset vector corresponding to the i-th detected text and the j-th template text is defined as (x2(i)-x1(j), y2(i)-y1(j)), and the i-th detected text box, the j-th template text box and the corresponding offset vector are added to the selected matching. Since the matched bill image has rigidity, and the perspective transformation can still maintain a certain locality, that is, the offset vectors corresponding to the text boxes with similar positions are similar, including the directionality of the vector and the length of the vector, and then the matching pairs are filtered. Take the similar matching pairs to perform box plot filtering on the direction and length to obtain a second reference configuration file.
[0081] Calculate the perspective transformation matrix of the ordered pair corresponding to the filtered reference template configuration file;
[0082] Based on the perspective transformation matrix and by using the random sample consensus algorithm (RANSAC), the second reference template configuration file is obtained through filtering and screening.
[0083] In one or more embodiments, the target template configuration file is determined according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box, comprising:
[0084] Determine the number of matches between each reference configuration file in the second reference configuration file and the text and image in the target bill image;
[0085] Perspective transform the template image corresponding to each reference configuration file in the second reference configuration file to the target bill image to obtain a perspective transformed image;
[0086] Determine the distance error between the matching points in the perspective transformed image; wherein the matching points include the center points of the text boxes in the target bill image and the center points of the text boxes corresponding to the template image;
[0087] Determine the target template configuration file according to the number of matches and the distance error.
[0088] Specifically, for example, the second reference configuration file with the most number of text matches in the target bill image and the smallest distance error of the matching points is configured as the target template configuration file.
[0089] In one or more embodiments, the bill information in the target bill image is extracted according to the target template configuration file, comprising:
[0090] Various types of information in the target bill image are extracted according to the target template configuration file.
[0091] Specifically, the extraction of the various types of information includes: extracting a fixed field. The labeled fixed field is projected and transformed using the calculated perspective transformation matrix, the nearest text box is found around, the nearest text box is taken as an extraction box, and the text content is taken as the extracted content. If the distance is greater than a threshold or the text boxes do not intersect, it is considered to be missed, and re-recognition can be performed at the text box position, and the text content is taken as the extracted content.
[0092] Linearly associated region information is extracted. The labeled linearly associated region is respectively projected and transformed using the calculated perspective transformation matrix to obtain a linear region. Further, the text boxes are screened to extract the text boxes intersecting with the linear region. The text lines are connected, specifically, the text boxes are selected from left to right and from top to bottom, a straight line is constructed with the center y coordinates of the text boxes to the right, and the text boxes passed by the straight line are regarded as the same text line. In this way, all the text lines are connected. According to the attributes of each region labeled, the attribute distribution of each text content of the text line is performed to obtain a list of all (key, val, unit) matching tuples.
[0093] Text block region information is extracted. The labeled text block region is projected and transformed using the calculated perspective transformation matrix, and the text blocks intersecting therewith are spliced to obtain text block content.
[0094] Image region information is extracted. The labeled text block region is projected and transformed using the calculated perspective transformation matrix, and the projected region is further processed according to the labeled process, such as image classification, segmentation, etc.
[0095] The error information in the extracted various types of information is corrected to obtain the bill information.
[0096] Specifically, the error correction methods include but are not limited to regular matching, similar word replacement, text correction, etc.
[0097] Date, numerical value type, etc. can be corrected by regular matching to filter out non-numeric strings; gender, age, etc. can be corrected by similar word replacement, and the characters "l" and "i" in age are replaced by the character "1"; text correction can be used in the presence of a dictionary, and the inspection items in a medical bill can be corrected to standard inspection items by the edit distance method.
[0098] Based on the above embodiment, in an application embodiment, as shown in Figure 5 The information extraction method of the bill image further includes:
[0099] Step S1: Investigate the bill invariant region and the region to be extracted, label the template file, and configure the template selection file.
[0100] Step S2: performing OCR text detection, text recognition;
[0101] Step S3: performing template matching;
[0102] Step S4: performing extraction of structured information;
[0103] Step S5: performing post-processing operation on the extracted fields.
[0104] As shown in Figure 6 Step S1 includes the following steps:
[0105] Step S11: performing template labeling. Mainly includes labeling positioning fields, labeling fixed fields, labeling linearly associated regions, labeling text block regions, etc. Except for positioning fields, the others are optional labeling, as shown in Figure 4 For illustration, the region where the pentagram is located is the positioning field, the region where the circle is located is the fixed field, and the region where the rectangle is located is the linearly associated region. The labeling types are explained as follows:
[0106] Positioning field: The purpose of the positioning field is to match the information extraction picture with the existing template. Since perspective transformation of the picture is involved, the positioning field must select fields with fixed relative positions and text contents, and at least four fields are included. The selected fields are preferably distributed in the largest possible area of the picture, such as the corners of the picture. The labeling method is as follows: the labeling type is selected as the positioning field, and the text content is the corresponding text content, obtaining a set of positioning fields {Xn} (n >= 4).
[0107] Fixed field: The fixed field is a field with fixed relative position and text content to be extracted, such as the name and gender fields on a medical test sheet. The labeling method is as follows: the labeling type is selected as the fixed field, and the labeling attribute key is the corresponding key to be recognized.
[0108] Linearly associated region: The linearly associated region is a plurality of associated regions, such as the test item column, test result column, and unit column on a medical test sheet. The labeling method is as follows: the labeling type is selected as the linearly associated region, and the labeling attributes correspond accordingly. For example, the item is labeled as key-1, the result is labeled as val-1, and the unit is labeled as unit-1. If there are two columns, the second column is labeled as key-2, val-2, and unit-2, and so on.
[0109] Text block region: The text block region is a whole block of text content to be recognized, such as medical order information on a medical bill. The labeling method is as follows: the labeling type is selected as the text block region, and the attribute key is the key to be recognized.
[0110] Picture region: picture region is the picture information to be processed, such as some medical photographs. The labeling method is: the labeling type is selected as a picture region, and the attribute is the processing method required by the picture and the allocated structured key.
[0111] Step S12: configuration of the template selection file. The template selection file is an optional item, and a tree structure is configured. In order to distinguish multiple templates, the fields contained in the template selection file are described as follows:
[0112] TextContain: followed by a text string T, indicating that if the text T is contained on the structured bill to be processed, the left child node is entered, and if the text T is not contained, the right child node is entered;
[0113] Output: followed by a template list [t1, t2,...tk], output node, indicating that the template list to be matched by the structured bill is [t1, t2,...tk];
[0114] Choice: (optional) followed by a selection method fun(ratio, err), indicating how to select the optimal matching template in the matching template.
[0115] As shown in Figure 7 , step S2 specifically includes the following steps:
[0116] Step S21: text box detection of the image, mainly including a detection-based method and a segmentation-based method, such as using a segmentation-based deep learning method DB-Net to perform text segmentation, and obtaining the BBox (Bounding Box) where the text is located through post-processing.
[0117] Step S22: text recognition, which is based on the text box detected in step S21 and performs text recognition on the detected text box. Mainly including CTC-Loss-based and Attention mechanism-based algorithms, such as using a CRNN deep learning network to perform text recognition on the detected BBox.
[0118] As shown in Figure 8 , step S3 specifically includes the following steps:
[0119] Step S31: preliminary screening of the template. The text box and text content obtained through step S2 are combined with the template selection configuration file of step S12 to perform preliminary filtering of the template. Specifically, the root node is entered, the selection of the subtree is performed according to the predicate TextContain, and the selection subtree is recursively performed, and finally the leaf node is entered, and the selected template set and the selection method of the optimal template are output.
[0120] Step S32: Template matching. The templates screened in step S31 are respectively matched with the picture to be structured and recognized. Specifically, the content of the positioning field labeled in step S11 is matched with the text content in step S2, and the text center coordinates and text content of the positioning field are recorded as: [(x1(1), y1(1), t1(1)), (x1(2), y1(2), t1(2),..., (x1(m), y1(m), t1(m))]. The text center coordinates and text content obtained through step S2 are: [(x2(1), y2(1), t2(1)), (x2(2), y2(2), t2(2),..., (x2(n), y2(n), t2(n))]. The Cartesian product of the template text and the detected text is traversed, if the text content of t2(i) contains the text content of t1(j), it is considered that the ith detected text and the jth template text are possibly matched, and the corresponding offset vector is defined as (x2(i)-x1(j), y2(i)-y1(j)), the ith detected text box and the jth template text box and the corresponding offset vector are added to the selected matching.
[0121] Step S33: Calculation of perspective transformation matrix. The matching pairs filtered through step S32 are calculated for the perspective transformation matrix, and the RANSAC algorithm is used for further filtering and screening.
[0122] Step S34: Final screening of templates. The matching number (n) and matching error (err) of the template to be recognized and each template are calculated according to steps S32 and S33. The matching error is defined as the distance error of the matching points after the template file is matched and perspective transformed to the to-be-recognized bill. The distance error can be selected as the Euclidean distance, and the greater the distance error, the greater the matching error. The number of positioning points of the template is recorded as N, and the matching rate ratio = n / N is defined. The optimal template is selected as follows:
[0123] 1. If the selection mode fun(ratio, err) is configured in the Choice field in step S12, the optimal template is selected according to the configured selection mode, that is, the matched template;
[0124] 2. Otherwise, the default mode is used: the templates with ratio > r0 (r0 is a defined constant) are filtered through threshold filtering, and then the matching error err is used as the evaluation standard in the templates meeting the threshold filtering condition, and the template with the minimum matching error is selected as the matched template.
[0125] As Figure 9 shown, step S4 comprises the following steps:
[0126] Step S41: Extract the fixed field. The fixed field marked by step S11 is projected and transformed using the perspective transformation matrix calculated in step S33, the nearest text box of step S2 is found around, the nearest text box is taken as the extraction box, and the text content is the extracted content. If the distance is greater than a threshold or the text boxes are not intersected, it is considered that the detection is missed, then re-recognition can be performed at the text box position, and the text content is taken as the extracted content.
[0127] Step S42: Extract the linearly associated region. The linearly associated region marked by step S11 is projected and transformed using the perspective transformation matrix calculated in step S33 to obtain a linear region. Further, the text boxes in step S2 are screened to extract the text boxes intersecting with the linear region. The text lines are connected, specifically, the text boxes are selected from left to right and from top to bottom, a straight line is constructed with the center y coordinates of the text boxes to the right, and the text boxes passed by the straight line are regarded as the same text line. In this way, all the text lines are connected. According to the properties of each region marked by step S11, the property distribution of each text content of the text line is performed, for example, as shown in the example of the linearly associated region in step S11, a list of all (key, val, unit) matching tuples is obtained.
[0128] Step S43: Extract the text block region. The text block region marked by step S11 is projected and transformed using the perspective transformation matrix calculated in step S33, and the text blocks intersecting therewith are spliced to obtain the text block content.
[0129] Step S44: Extract the image region. The text block region marked by step S11 is projected and transformed using the perspective transformation matrix calculated in step S33, and the projected region is further processed according to the steps marked by S11, such as image classification, segmentation, etc.
[0130] Step S5 is post-processing, including regular matching, similar word replacement, text error correction, etc. Date, numerical type, etc. can use regular matching to correct errors and filter out non-numeric strings;
[0131] Gender, age, etc. can use similar word replacement to correct errors, and the characters “l” and “i” in age are replaced with the character “1”.
[0132] Text error correction can be used in the presence of a dictionary, and the inspection items in a medical bill can be corrected to standard inspection items by using the edit distance method.
[0133] The embodiments of the present application also have the following beneficial effects:
[0134] Compared to manual entry of paper bills, the automated information extraction process of paper bills in this embodiment of the present invention can save a significant amount of labor while achieving higher accuracy. The real-time nature of electronic, structured data acquisition can also be significantly improved, allowing for timely subsequent application processes. Compared to rule-based methods, template-matching-based information extraction offers the advantage of single-step template annotation and universal use of identical templates, eliminating the need for manual rule-making based on location and other information.
[0135] Compared with existing template matching methods, this method adds a template configuration process. Template selection can reduce the number of matching templates and distinguish similar templates, improving real-time performance and accuracy. Template positioning field annotation supports annotation of identical text fields and has a certain degree of fault tolerance for annotation errors, improving fault tolerance and robustness. It can also support sub-field types with template-matching content, adding linear, text block, image and other areas, enhancing the versatility and scalability of template matching.
[0136] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0137] According to another aspect of the embodiment of the present invention, there is also provided a bill image information extraction device for implementing the above-mentioned bill image information extraction method. Figure 10 As shown, the device includes:
[0138] The recognition unit 1002 recognizes the target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes;
[0139] An acquiring unit 1004 is configured to acquire a first reference template configuration file from a template configuration file set based on the text information in each of the identified text boxes; wherein each template configuration file in the template configuration file set is a template file obtained by annotating fields based on a document image, and the template configuration file set includes multiple predicates associated via a tree structure;
[0140] A screening unit 1006 is configured to screen the first reference template configuration file to obtain a second reference template configuration file; wherein the text information of the positioning field in the template configuration file in the second reference template configuration file includes the text information in the identification text box;
[0141] The determining unit 1008 is configured to determine a target template configuration file according to a matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box.
[0142] The extracting unit 1010 is configured to extract the bill information in the target bill image according to the target template configuration file.
[0143] In the embodiment of the present application, the target bill image is identified to obtain a set of identified text boxes and text information in each identified text box; a first reference template configuration file is obtained from a set of template configuration files based on the text information in each identified text box; each template configuration file in the set of template configuration files is a template file obtained by annotating types and fields of bill images, and the set of template configuration files includes multiple predicates associated by a tree structure; the first reference template configuration file is filtered to obtain a second reference template configuration file; the text information of the positioning field in the template configuration file in the second reference template configuration file contains the text information in the identified text box; a target template configuration file is determined according to a matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box; and the bill information in the target bill image is extracted according to the target template configuration file. In the above method, since a target bill template with a higher matching degree is filtered from the annotated types and regions of the bill template, and then the information in the bill image is extracted according to the target bill template, not only the information in the bill image can be accurately obtained, but also the efficiency of extracting information of paper bills can be improved, thereby solving the technical problem of low efficiency of extracting information of paper bills.
[0144] In one or more embodiments, the identifying unit 1002 specifically includes:
[0145] The segmentation module is configured to process the target bill image based on an image segmentation model to obtain the set of identified text boxes.
[0146] The recognition module is configured to identify the set of text boxes by using a character recognition model to obtain the text information in each identified text box.
[0147] In one or more embodiments, the bill image information extraction apparatus further includes:
[0148] The annotation unit is configured to perform type annotation on regions where various types of information in a bill template picture are located.
[0149] The dividing unit is configured to divide the labeled bill template picture into a tree structure according to text content in the bill template picture, and obtain the template configuration file set including multiple predicates connected by the tree structure.
[0150] In one or more embodiments, the labeling unit includes at least one of:
[0151] The first labeling module is configured to label a positioning field in the bill template picture to obtain a labeled positioning field set, wherein the positioning field is a field with fixed relative position and text content in the bill template picture.
[0152] The second labeling module is configured to label a fixed field in the bill template picture to obtain a labeled fixed field set, wherein the fixed field is a field with fixed relative position in the bill template picture.
[0153] The third labeling module is configured to label a linearly associated tuple in the bill template picture to obtain a labeled linearly associated tuple set, wherein the linearly associated tuple includes multiple associated fields.
[0154] The fourth labeling module is configured to label a text block region in the bill template picture to obtain a labeled text block region set.
[0155] The fifth labeling module is configured to label an image region in the bill template picture to obtain a labeled image set.
[0156] In one or more embodiments, the obtaining unit 1004 specifically includes:
[0157] The first determining module is configured to determine an initial matching field according to text information in all recognized text boxes.
[0158] The traversal module is configured to take the initial matching field as a root node, take the remaining fields of all recognized text boxes as child nodes in turn, traverse each multiple predicate in the template configuration file set, and obtain the first reference template configuration file.
[0159] In one or more embodiments, the screening unit 1006 includes:
[0160] The first obtaining module is configured to obtain a positioning field set included in all template configuration files in the first reference template configuration file.
[0161] The second determining module is configured to determine a Cartesian product of text information in each recognized text box and the positioning field set in the first reference template configuration file.
[0162] The screening module is configured to screen the second reference template configuration file from the first reference template configuration file according to a reference template configuration file corresponding to the ordered pair text in the Cartesian product.
[0163] In one or more embodiments, the screening module comprises:
[0164] The first determination subunit is configured to determine an offset vector corresponding to the text boxes of the two texts of the ordered pair with the same text.
[0165] The first filtering subunit is configured to filter through a box plot according to the directionality and the modulus length attribute of the offset vector to obtain a filtered reference template configuration file.
[0166] The calculation subunit is configured to calculate a perspective transformation matrix of the ordered pair corresponding to the filtered reference template configuration file.
[0167] The second filtering subunit is configured to filter and screen based on the perspective transformation matrix and through a random sample consensus algorithm (RANSAC) to obtain the second reference template configuration file.
[0168] In one or more embodiments, the determination unit 1008 comprises:
[0169] The third determination module is configured to determine the number of matches of each reference configuration file in the second reference configuration file with the text and the image in the target bill image.
[0170] The perspective transformation module is configured to perform perspective transformation on the template image corresponding to each reference configuration file in the second reference configuration file to the target bill image to obtain a perspective transformation image.
[0171] The fourth determination module is configured to determine the distance error between the matching points in the perspective transformation image, wherein the matching points include the center points of the text boxes in the target bill image and the center points of the text boxes corresponding to the template image.
[0172] The fifth determination module is configured to determine the target template configuration file according to the number of matches and the distance error.
[0173] In one or more embodiments, the extraction unit 1010 comprises:
[0174] The extraction module is configured to extract various types of information in the target bill image according to the target template configuration file.
[0175] The error correction module is configured to correct the error information in the extracted various types of information to obtain the bill information.
[0176] According to a further aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above information extraction method of a bill image, which can be Figure 11 The electronic device can be a terminal device or a server, as shown in the figure. The present embodiment takes the terminal device as an example for illustration. As shown in the figure, the electronic device comprises a memory 1102 and a processor 1104, wherein the memory 1102 stores a computer program, and the processor 1104 is configured to execute the steps in any of the above method embodiments through the computer program. Figure 11
[0177] Optionally, in the present embodiment, the electronic device can be located in at least one of a plurality of network devices in a computer network.
[0178] Optionally, in the present embodiment, the processor can be configured to execute the following steps through the computer program:
[0179] S1, identifying a target bill image to obtain a set of identified text boxes and text information in each identified text box in the set of identified text boxes;
[0180] S2, obtaining a first reference template configuration file from a set of template configuration files based on the text information in each identified text box; wherein each template configuration file in the set of template configuration files is a template file obtained by annotating the types and fields of bill images, and the set of template configuration files comprises a plurality of predicates associated through a tree structure;
[0181] S3, screening the first reference template configuration file to obtain a second reference template configuration file; wherein the text information of the positioning field in the template configuration file in the second reference template configuration file contains the text information in the identified text box;
[0182] S4, determining a target template configuration file according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box;
[0183] S5, extracting bill information in the target bill image according to the target template configuration file.
[0184] Optionally, those skilled in the art can understand that Figure 11 The structure shown in the figure is only schematic, and the electronic device can also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, etc. Figure 11 It does not limit the structure of the electronic device. For example, the electronic device can further include more or less components (such as a network interface, etc.) than those shown in Figure 11 or have a different configuration than that shown in Figure 11 .
[0185] The memory 1102 can be used to store software programs and modules, such as program instructions / modules corresponding to the information extraction method and device of the bill image in the embodiments of the present application. The processor 1104 executes various functions and data processing by running the software programs and modules stored in the memory 1102, that is, implements the information extraction method of the bill image as described above. The memory 1102 can include a high-speed random access memory, and can further include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 1102 can further include a memory remotely arranged with respect to the processor 1104, which can be connected to the terminal through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. Specifically, the memory 1102 can be but is not limited to used for storing information such as information of multiple bill images. As an example, as shown in Figure 11 the memory 1102 can include but is not limited to the identification unit 1002, the acquisition unit 1004, the screening unit 1006, the determination unit 1008, and the extraction unit 1010 in the information extraction device of the bill image. In addition, other module units in the information extraction device of the bill image can also be included, which will not be described herein.
[0186] Optionally, the transmission device 1106 is configured to receive or send data via a network. Specific examples of the network can include wired networks and wireless networks. In one example, the transmission device 1106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable to communicate with the Internet or a local area network. In one example, the transmission device 1106 is a radio frequency (Radio Frequency, RF) module, which is configured to communicate with the Internet in a wireless manner.
[0187] In addition, the electronic device further includes a display 1108 configured to display the identification result of the bill image, and a connection bus 1110 configured to connect various module components in the electronic device.
[0188] In other embodiments, the terminal device or the server described above can be a node in a distributed system, where the distributed system can be a blockchain system, which can be a distributed system formed by the plurality of nodes connected through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, a terminal, or an electronic device, can become a node in the blockchain system by joining the P2P network.
[0189] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the information extraction method of the bill image described above, where the computer program is configured to execute the steps in any of the method embodiments described above when running.
[0190] Optionally, in the present embodiment, the computer readable storage medium described above can be configured to store a computer program for executing the following steps:
[0191] S1, identifying a target bill image to obtain a set of identified text boxes and text information in each identified text box in the set of identified text boxes;
[0192] S2, obtaining a first reference template configuration file from a set of template configuration files based on the text information in each identified text box, where each template configuration file in the set of template configuration files is a template file obtained by annotating the types and fields of bill images, and the set of template configuration files includes multiple predicates associated through a tree structure;
[0193] S3, screening the first reference template configuration file to obtain a second reference template configuration file, where the text information of the positioning field in the template configuration file in the second reference template configuration file contains the text information in the identified text box;
[0194] S4, determining a target template configuration file according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box;
[0195] S5, extracting bill information in the target bill image according to the target template configuration file.
[0196] Optionally, in the embodiment, a person skilled in the art can understand that all or part of the steps of the various methods in the above embodiment can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, and the storage medium can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0197] The serial numbers of the above embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0198] The integrated units in the above embodiments, if realized in the form of software function units and sold or used as independent products, can be stored in the above computer readable storage medium. Based on such understanding, the technical solutions of the application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the embodiments of the application.
[0199] In the above embodiments of the application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0200] In several embodiments provided in the present application, it should be understood that the disclosed client can be implemented by other ways. Among them, the above described device embodiments are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, and can be electrical or other forms.
[0201] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0202] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0203] The above is only the preferred embodiment of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. A method of information extraction of a document image, characterized by, The method comprises the following steps: recognizing a target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes; obtaining a first reference template configuration file from a set of template configuration files based on the text information in each recognized text box, wherein each template configuration file in the set of template configuration files is a template file obtained by annotating types and fields according to a bill image, and the set of template configuration files comprises multiple predicates associated by a tree structure; screening the first reference template configuration file to obtain a second reference template configuration file, which comprises the following steps: obtaining a set of positioning fields contained in all template configuration files in the first reference template configuration file; determining a Cartesian product of the text information in each recognized text box and the set of positioning fields in the first reference template configuration file; determining offset vectors corresponding to text boxes where two texts of an ordered pair with the same text in the Cartesian product are located according to the reference template configuration file corresponding to the ordered pair with the same text; filtering the offset vectors according to their directionality and modulus to obtain a filtered reference template configuration file by means of a box plot; calculating a perspective transformation matrix of the ordered pair corresponding to the filtered reference template configuration file; and filtering and screening the perspective transformation matrix by means of a random sample consensus algorithm to obtain the second reference template configuration file; wherein the text information of the positioning fields in the template configuration file in the second reference template configuration file contains the text information in the recognized text box; determining a target template configuration file according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each recognized text box; extracting bill information in the target bill image according to the target template configuration file.
2. The method of claim 1, wherein, The method of recognizing a target bill image to obtain a set of recognized text boxes and text information in each recognized text box in the set of recognized text boxes comprises the following steps: processing the target bill image based on an image segmentation model to obtain the set of recognized text boxes; recognizing the set of text boxes by means of a character recognition model to obtain the text information in each recognized text box.
3. The method of claim 1, wherein, Before the target bill image is recognized, the method further comprises the following steps: annotating the types of regions where various types of information in a bill template image are located; dividing the annotated bill template image into the set of template configuration files comprising multiple predicates associated by a tree structure by taking a preset word as a root node and according to the text content in the bill template image.
4. The method of claim 3, wherein, The step of annotating the types of regions where various types of information in a bill template image are located comprises at least one of the following steps: annotating positioning fields in the bill template image to obtain a set of annotated positioning fields, wherein the positioning fields are fields with fixed relative positions and text content in the bill template image; annotating fixed fields in the bill template image to obtain a set of annotated fixed fields; wherein the fixed fields are fields with fixed relative positions in the bill template image. Annotate the linearly associated tuples in the bill template picture to obtain an annotated linearly associated tuple set; wherein the linearly associated tuples include a plurality of associated fields; Annotate the text block regions in the bill template picture to obtain an annotated text block region set; Annotate the image regions in the bill template picture to obtain an annotated image set.
5. The method of claim 1, wherein, The first reference template configuration file is obtained from a template configuration file set based on the text information in each identified text box, including: Determine the initial matching field according to the text information in all identified text boxes; Take the initial matching field as a root node, and take the remaining fields of all identified text boxes as child nodes in turn, traverse each multi-predicate in the template configuration file set, and obtain the first reference template configuration file.
6. The method of claim 1, wherein, The second reference template configuration file is obtained based on the perspective transformation matrix and filtered and screened through the random sample consensus algorithm, including: The second reference template configuration file is obtained based on the perspective transformation matrix and filtered and screened through the random sample consensus algorithm RANSAC.
7. The method of claim 1, wherein, The target template configuration file is determined according to the matching degree of each reference configuration file in the second reference configuration file and the text information in each identified text box, including: Determine the matching number of each reference configuration file in the second reference configuration file and the text and image in the target bill image; Perspective transform the template image corresponding to each reference configuration file in the second reference configuration file to the target bill image to obtain a perspective transformation image; Determine the distance error between matching points in the perspective transformation image; wherein the matching points include the center points of the text boxes in the target bill image and the center points of the text boxes corresponding to the template image; Determine the target template configuration file according to the matching number and the distance error.
8. The method of claim 1, wherein, The bill information in the target bill image is extracted according to the target template configuration file, including: Various types of information in the target bill image are extracted according to the target template configuration file; Error correction is performed on the error information in the extracted various types of information to obtain the bill information.
9. A device for extracting information from bill images, characterized in that: Including: An identification unit identifies a target bill image to obtain an identified text box set and text information in each identified text box in the identified text box set; An acquisition unit is configured to obtain a first reference template configuration file from a template configuration file set based on text information in each identified text box; wherein each template configuration file in the template configuration file set is a template file obtained by annotating fields according to a bill image, and the template configuration file set includes multi-predicates associated through a tree structure; The screening unit is configured to screen the first reference template configuration file to obtain a second reference template configuration file, and includes: obtaining a set of positioning fields contained in all template configuration files in the first reference template configuration file; determining a Cartesian product of text information in each identified text box and the set of positioning fields in the first reference template configuration file; determining, according to a reference template configuration file corresponding to an ordered pair of texts with the same text in the Cartesian product, an offset vector corresponding to text boxes of the two texts of the ordered pair of texts with the same text; filtering the offset vector according to directionality and a modulus length attribute of the offset vector by using a box plot to obtain a filtered reference template configuration file; calculating a perspective transformation matrix of an ordered pair corresponding to the filtered reference template configuration file; and filtering and screening, based on the perspective transformation matrix and by using a random sample consensus algorithm, to obtain the second reference template configuration file; wherein text information of the positioning fields in a template configuration file in the second reference template configuration file contains text information in an identified text box. The determining unit is configured to determine a target template configuration file according to a matching degree of each reference configuration file in the second reference configuration file and text information in each identified text box. The extracting unit is configured to extract bill information in the target bill image according to the target template configuration file.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method in any one of claims 1 to 8 by using the computer program.
11. A computer readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, and the program is configured to execute the method in any one of claims 1 to 8 when running.
Citation Information
Patent Citations
Bill image recognition method and device, electronic equipment and storage medium
CN112669515A
Bill classification method, bill classification device, electronic equipment and storage medium
CN114140649A