A method and apparatus for extracting text from a bill image
Through the text detection and recognition model and clustering algorithm, the fields in the bill image are automatically identified and typed, which solves the problem of manual adjustment of text and layout formats in the prior art, and improves the efficiency of medical bill processing.
Patent Information
- Application Number
- CN202210277776.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-03-21
AI Technical Summary
The existing technology cannot automatically adjust the identified text and the format of the bill layout, resulting in inefficient entry and review of medical bills and requires a lot of manual intervention.
The text detection and recognition model recognizes the coordinates of the field center point in the bill image, uses the clustering algorithm to determine the rows and columns of the fields, and automatically typesets according to the coordinate difference value, and outputs text that meets the bill layout.
Automatic layout of bill image text is realized, reducing labor losses, improving efficiency, and reducing the need for manual adjustment.
Smart Images

Figure CN114612922B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bill recognition, and particularly to a method and device for extracting text from a bill image. Background Art
[0002] In recent years, the medical industry in China has been continuously developing and progressing, and the system and regulations of the medical industry tend to be mature and perfect. However, the formats of bills issued by each hospital are different. After the medical bills are submitted, the reimbursement unit enters the information manually, the verification personnel verify the entered information, and then submit it to the review personnel for review. The entire process has low work efficiency, requires a large amount of human resources, and there will also be some errors that are difficult to detect.
[0003] In response to this, the existing technology uses a text detection and recognition model to recognize the text of each medical bill. However, the current technology can only simply recognize the text in the bill and cannot typeset the recognized text according to the layout format of different bills. Usually, it is necessary to manually adjust and typeset the recognized text to make it consistent with the layout format of the bill. Therefore, the problem of low work efficiency still exists. Summary of the Invention
[0004] The embodiments of the present invention provide a method and device for extracting text from a bill image, which can automatically recognize the text of the bill image and automatically typeset the text according to the layout format of the bill and then output it, reducing human consumption and improving efficiency.
[0005] An embodiment of the present invention provides a method for extracting text from a bill image, including: obtaining a bill image, inputting the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image;
[0006] Clustering each field according to the difference in the ordinate of the center point of each field, and then taking the fields in the same category as the fields in the same row; determining the row order of each row according to the average value of the ordinate of the center point of the fields in each row; determining the sorting of each field in the corresponding row according to the abscissa of the center point of the field;
[0007] Clustering each field according to the difference in the abscissa of the center point of each field, and then taking the fields in the same category as the fields in the same column; determining the column order of each column according to the average value of the abscissa of the center point of the fields in each column; determining the sorting of each field in the corresponding column according to the ordinate of the center point of the field;
[0008] Performing typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and outputting each field according to the layout format after typesetting.
[0009] Further, before inputting the bill image into a preset text detection and recognition model, it further includes: performing image preprocessing on the bill image; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the bill angle, and removing the bill official seal.
[0010] Further, the text detection and recognition model includes: a text detection sub-model and a text recognition model; the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image, specifically including:
[0011] Detecting the bill image through the text detection sub-model, and recognizing the rectangular frame coordinates of each suspected field and the confidence score corresponding to each suspected field;
[0012] Removing the rectangular frames of the suspected fields with confidence scores less than a preset threshold, then intercepting the corresponding rectangular frame images according to the rectangular frame coordinates of the remaining suspected fields, and inputting the intercepted rectangular frame images into the text recognition sub-model, so that the text recognition sub-model recognizes the fields, the abscissa of the center line of the fields, and the ordinate of the center line of the fields in each rectangular frame image.
[0013] Further, after determining the sorting of each field in the corresponding row according to the abscissa of the center point of the field, it further includes:
[0014] Clustering the fields in the same row based on the clustering algorithm according to the difference in the abscissa of the center points of adjacent fields in the same row, and taking the category with the largest number of fields as the first reference category;
[0015] Calculating the average value of the differences in the abscissa of the center points in the first reference category to obtain a first average value, and taking the first average value as the column width of the corresponding row;
[0016] Judging one by one whether the difference in the abscissa of the center points of each adjacent field is greater than the column width of the corresponding row. If so, filling a first preset character between the two adjacent fields.
[0017] Further, after determining the sorting of each field in the corresponding column according to the ordinate of the center point of the field, it further includes:
[0018] Clustering the fields in the same column based on the clustering algorithm according to the difference in the ordinate of the center points of adjacent fields in the same column, and taking the category with the largest number of fields as the second reference category;
[0019] Calculating the average value of the differences in the ordinate of the center points in the second reference category to obtain a second average value, and taking the second average value as the row height of the corresponding column;
[0020] Judging one by one whether the difference in the ordinate of the center points of adjacent fields is greater than the line height of the corresponding column. If so, filling a second preset character between the two adjacent fields.
[0021] Further, typesetting is performed according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and each field is output according to the typeset page format, including:
[0022] Typesetting is performed according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, the sorting of each field in the corresponding column, the first preset filling character in each row, and the second preset filling character in each column, and each field, each first preset filling field, and each second preset filling character are output according to the typeset page format.
[0023] Further, each field, each first preset filling field, and each second preset filling character are output according to the typeset page format, including:
[0024] Each field, each first preset filling field, and each second preset filling character are output in json format or excel file format according to the typeset page format.
[0025] Based on the above method embodiments, the present invention correspondingly provides a text extraction device for a bill image, including: a field recognition module, a row determination module, a column determination module, and a typesetting module;
[0026] The field recognition module is used to obtain a bill image and input the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image;
[0027] The row determination module is used to cluster each field according to the difference in the ordinate of the center points of each field, and then take the fields in the same class as the fields in the same row; determine the row order of each row according to the average value of the ordinate of the center points of the fields in each row; determine the sorting of each field in the corresponding row according to the abscissa of the center point of the field;
[0028] The column determination module is used to cluster each field according to the difference in the abscissa of the center points of each field, and then take the fields in the same class as the fields in the same column; determine the column order of each column according to the average value of the abscissa of the center points of the fields in each column; determine the sorting of each field in the corresponding column according to the ordinate of the center point of the field;
[0029] The typesetting module is used to perform typesetting according to the line order of each line, the column order of each column, the sorting of each field in the corresponding line, and the sorting of each field in the corresponding column, and output each field according to the typeset page format.
[0030] Further, it further includes: an image preprocessing module, which is used to perform image preprocessing on the bill image before the field recognition module inputs the bill image into a preset text detection and recognition model; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the angle of the bill, and removing the official seal of the bill.
[0031] By implementing the embodiments of the present invention, the following beneficial effects are achieved:
[0032] The embodiments of the present invention provide a method and device for extracting text from a bill image. After the method recognizes the text in the bill image through a text detection and recognition model, it determines the rows and columns where each field is located, and the order of each field in the corresponding row and column through a clustering algorithm. Finally, it automatically typesets and outputs the recognized text according to the determined rows and columns where each field is located, and the order of each field in the corresponding row and column, realizing automatic typesetting of fields, eliminating the need for manual typesetting, reducing labor consumption and improving efficiency. Description of the Drawings
[0033] Figure 1 is a schematic flowchart of a method for extracting text from a bill image provided by an embodiment of the present invention.
[0034] Figure 2 is a schematic structural diagram of a device for extracting text from a bill image provided by an embodiment of the present invention. Detailed Embodiments
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0036] As Figure 1 shown, an embodiment of the present invention provides a method for extracting text from a bill image, including:
[0037] Step S101: Obtain a bill image, and input the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image;
[0038] Step S102: Cluster each field according to the difference in the ordinate of the center point of each field, and then take the fields in the same class as the fields in the same row; determine the row order of each row according to the average value of the ordinate of the center point of the fields in each row; determine the sorting of each field in the corresponding row according to the abscissa of the center point of the field.
[0039] Step S103: Cluster each field according to the difference in the abscissa of the center point of each field, and then take the fields in the same class as the fields in the same column; determine the column order of each column according to the average value of the abscissa of the center point of the fields in each column; determine the sorting of each field in the corresponding column according to the ordinate of the center point of the field.
[0040] Step S104: Perform typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and output each field according to the layout format after typesetting.
[0041] For step S101, in a preferred embodiment, before inputting the bill image into a preset text detection and recognition model, it further includes: performing image preprocessing on the bill image; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the bill angle, and removing the bill official seal.
[0042] Specifically, in this embodiment of the present invention, the bill image is obtained by scanning a document, and then the opencv image processing library is used to perform image preprocessing on the bill image. The main image preprocessing includes:
[0043] ① Remove blurred images by detecting image blurriness: Blurred images have less boundary information, while normal images have more clear boundary information. Therefore, the variance of the second derivative of the image can be used as the basis for determining whether the image is blurred; based on the existing Laplacian operator having second-order differentiability, the area (boundary) with rapid density change in the bill image can be calculated. Therefore, calculate the Laplacian operator for the bill image and calculate the variance. The smaller the variance, the more blurred the image. The blur threshold can be dynamically set according to the detection needs. If the variance value of the bill image is less than this blur threshold, it is determined that the bill image is blurred and the bill image is removed.
[0044] ②Realize bill angle correction by correcting the horizontal state of the image: In the process of obtaining the bill image by scanning, it is necessary to manually place the paper bill, which has a certain angular deviation, and the obtained bill image is inclined to a certain extent. Therefore, it is necessary to perform bill angle correction. Specifically, in each bill, there is one or more straight lines. Use the morphological functions erode and dilate of the opencv image processing library to find the straight lines. Since there may be multiple straight lines, and only one straight line is needed to calculate the inclination angle, use the findcontours function of opencv to calculate the straight line with the largest perimeter and return the angle of this straight line. Finally, use the warpAffine function of opencv to rotate the image to make the bill image in a horizontal state. This method can process bills at any angle.
[0045] ③Remove the red official seal from the bill to eliminate the official seal of the bill: Generally, the bill is stamped with the official seal of the issuing unit, which will affect the accuracy of recognizing the text. Use the split function of opencv to separate the image channels to obtain the R, G, and B images. Then subtract the R image from the G image. After subtraction, perform binaryzation processing on the bill image. After binaryzation, obtain the coordinates of the white pixel points in the bill image, that is, the coordinates of the red seal position. Finally, use white pixels to mask the red pixels on the original bill image.
[0046] In a preferred embodiment, the text detection and recognition model includes: a text detection sub-model and a text recognition model;
[0047] The text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image, specifically including:
[0048] Detect the bill image through the text detection sub-model and recognize the rectangular box coordinates of each suspected field and the confidence score corresponding to each suspected field;
[0049] Remove the rectangular boxes of the suspected fields with confidence scores less than the preset threshold. Then, intercept the corresponding rectangular box images according to the rectangular box coordinates of the remaining suspected fields, and input the intercepted rectangular box images into the text recognition sub-model so that the text recognition sub-model recognizes the fields, the abscissa of the center line of the fields, and the ordinate of the center line of the fields in each rectangular box image.
[0050] In the present invention, the text detection sub-model can use the DBnet text detection network to output the rectangular box coordinates and confidence scores of the fields; by setting a preset threshold, the text rectangular boxes with confidence scores less than the preset threshold are eliminated; finally, according to the remaining text rectangular boxes, the text rectangular box images are intercepted and input into the text recognition sub-model constructed by CNN+RNN+CTC, and the corresponding fields, the abscissa and ordinate of the center point of the field in the text rectangular box image are obtained;
[0051] For step S102: When clustering, the K-Means clustering algorithm is adopted. Calculate the difference in the ordinate of the center point between each pair of fields one by one, and then use the K-Means clustering algorithm for clustering. If two fields are in the same row, then the difference is relatively small. If two fields are in different rows, then the difference in the ordinate will be relatively large. Based on this principle, the K-Means clustering algorithm can cluster each field, so that each field can be divided into several rows. After dividing each field into the corresponding row, calculate the average value of the ordinate of the center point of all fields in this row, and determine the row order of each row according to the average value. In this way, it can be determined which specific row each field is in. Then, determine the sorting of each field in the corresponding row according to the abscissa of the center point of the field, so that it can be determined which position each field is in the corresponding row.
[0052] In a preferred embodiment, after determining the sorting of each field in the corresponding row according to the abscissa of the center point of the field, it further includes: Based on the difference in the abscissa of the center point between adjacent fields in the same row, cluster the fields in the same row based on the clustering algorithm, and use the category with the largest number of fields as the first reference category; calculate the average value of the differences in the abscissa of the center points in the first reference category to obtain the first average value, and use the first average value as the column width of the corresponding row; judge one by one whether the difference in the abscissa of the center point between each adjacent field is greater than the column width of the corresponding row. If so, fill the first preset character between the two adjacent fields.
[0053] According to the above steps, the specific rows where each field is located and their sorting in each row have been determined. However, there is a problem here. If a field in a certain column is missing in the same row, after sorting from smallest to largest, the subsequent fields will fill forward, and ultimately the layout content will be disordered during subsequent typesetting. Therefore, in this embodiment, the column width corresponding to each row is further calculated through a clustering algorithm. Based on the calculated column width, for the fields sorted in the same row, the difference in the abscissa of the center points between adjacent fields is calculated. If the difference exceeds the column width, it is determined that there is a missing field between the two fields. At this time, the first preset character can be used to fill between the two fields. Optionally, the above first preset field can be a double quotation mark or a space character, etc., which can be adjusted according to the actual situation. Through this embodiment, the correctness of subsequent typesetting can be maintained even when there are missing fields.
[0054] For step S103, calculate the difference in the abscissa of the center points between each pair of fields one by one, and then use the K-Means clustering algorithm for clustering. If two fields are in the same column, the difference is relatively small. If two fields are in different columns, the difference in the abscissa is relatively large. Based on this principle, the K-Means clustering algorithm can cluster the fields, so that the fields can be divided into several columns. After dividing the fields into the corresponding columns, calculate the average value of the abscissa of the center points of all fields in this column. Determine the column order of each row based on the average value, so that it can be determined in which specific column each field is located. Then, determine the sorting of each field in the corresponding column according to the ordinate of the center point of the field, so that it can be determined the position of each field in the corresponding column.
[0055] In a preferred embodiment, after determining the sorting of each field in the corresponding column according to the ordinate of the center point of the field, it further includes:
[0056] Based on the difference in the ordinate of the center points of adjacent fields in the same column, cluster the fields in the same column using the clustering algorithm, and take the category with the largest number of fields as the second reference category;
[0057] Calculate the average value of the differences in the ordinate of the center points in the second reference category to obtain a second average value, and use the second average value as the line height of the corresponding column;
[0058] Judge one by one whether the difference in the ordinate of the center points of adjacent fields is greater than the line height of the corresponding column. If so, fill a second preset character between the two adjacent fields.
[0059] According to the above steps, the specific columns where each field is located and the sorting in each column have been determined. However, there is a problem here. If a field in a certain column is missing, after sorting from smallest to largest, the subsequent fields will fill forward. Eventually, when typesetting later, the content on the page will be disordered. Therefore, in this embodiment, the row height corresponding to each column is further calculated through a clustering algorithm. According to the calculated row height, for the fields sorted in the same column, calculate the difference in the vertical coordinates of the centers between adjacent fields. If the difference exceeds the row height, then it is determined that there is a missing field between the two fields. At this time, a second preset character can be used to fill between the two fields; optionally, the above second preset field can be a double quotation mark or a space character, etc., and can be specifically adjusted according to the actual situation. Through this embodiment, it is possible to maintain the correctness of subsequent typesetting even when there are missing fields.
[0060] For step S104, if after recognition, there are no missing fields in each row and each column, then at this time, the fields can be directly typeset according to the row sequence of each row, the column sequence of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and the fields are output in json format or excel file format according to the typeset page format.
[0061] If after recognition, there are missing fields in each row and each column, then in a preferred embodiment, typesetting is performed according to the row sequence of each row, the column sequence of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and the fields are output according to the typeset page format, including:
[0062] Typesetting is performed according to the row sequence of each row, the column sequence of each column, the sorting of each field in the corresponding row, the sorting of each field in the corresponding column, the first preset filling character in each row, and the second preset filling character in each column, and the fields, each first preset filling field, and each second preset filling character are output according to the typeset page format. Among them, outputting the fields, each first preset filling field, and each second preset filling character according to the typeset page format includes: outputting the fields, each first preset filling field, and each second preset filling character in json format or excel file format according to the typeset page format.
[0063] Based on the above method item embodiment, the present invention correspondingly provides a device item embodiment;
[0064] As Figure 2 shown, an embodiment of the present invention provides a text extraction device for a bill image, including: a field recognition module 1, a row determination module 2, a column determination module 3, and a typesetting module 4;
[0065] The field recognition module is configured to obtain a bill image and input the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image;
[0066] The row determination module is configured to cluster each field according to the difference in the ordinate of the center point of each field, and then regard the fields in the same class as the fields in the same row; determine the row order of each row according to the average value of the ordinate of the center point of the fields in each row; determine the sorting of each field in the corresponding row according to the abscissa of the center point of the field;
[0067] The column determination module is configured to cluster each field according to the difference in the abscissa of the center point of each field, and then regard the fields in the same class as the fields in the same column; determine the column order of each column according to the average value of the abscissa of the center point of the fields in each column; determine the sorting of each field in the corresponding column according to the ordinate of the center point of the field;
[0068] The typesetting module is configured to perform typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and output each field according to the typeset page format.
[0069] In a preferred embodiment, it further includes: an image preprocessing module, which is configured to perform image preprocessing on the bill image before the field recognition module inputs the bill image into a preset text detection and recognition model; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the bill angle, and removing the bill official seal.
[0070] It should be noted that the device embodiment of the present invention corresponds to the method embodiment of the present invention, and it can implement the text extraction method of the bill image described in any one of the present invention. In addition, the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiment provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative labor.
[0071] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications are also regarded as the protection scope of the present invention.
Claims
1. A method for extracting text from a bill image, characterized in that, Including: Obtain a bill image, and input the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image; Cluster each field according to the difference in the ordinate of the center point of each field, and then regard the fields in the same class as the fields in the same row; determine the row order of each row according to the average value of the ordinate of the center point of the fields in each row; Determine the sorting of each field in the corresponding row according to the abscissa of the center point of the field; Cluster the fields in the same row based on the clustering algorithm according to the difference in the abscissa of the center point of adjacent fields in the same row, and regard the class with the largest number of fields as the first reference class; Calculate the average value of the differences in the abscissa of the center points in the first reference class to obtain a first average value, and use the first average value as the column width of the corresponding row; judge one by one whether the difference in the abscissa of the center points of adjacent fields is greater than the column width of the corresponding row. If so, fill a first preset character between the two adjacent fields; Cluster each field according to the difference in the abscissa of the center point of each field, and then regard the fields in the same class as the fields in the same column; determine the column order of each column according to the average value of the abscissa of the center point of the fields in each column; Determine the sorting of each field in the corresponding column according to the ordinate of the center point of the field; Perform typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and output each field according to the typeset page format.
2. The method for extracting text from a bill image according to claim 1, wherein Before inputting the bill image into the preset text detection and recognition model, it further includes: Perform image preprocessing on the bill image; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the bill angle, and removing the bill official seal.
3. The method for extracting text from a bill image according to claim 1, wherein: The text detection and recognition model includes: a text detection sub-model and a text recognition sub-model; The text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image, specifically including: Detect the bill image through the text detection sub-model, and recognize the rectangular frame coordinates of each suspected field and the confidence score corresponding to each suspected field; Remove the rectangular frames of the suspected fields with confidence scores less than the preset threshold, and then intercept the corresponding rectangular frame images according to the rectangular frame coordinates of the remaining suspected fields, and input the intercepted rectangular frame images into the text recognition sub-model, so that the text recognition sub-model recognizes the fields, the abscissa of the center line of the field, and the ordinate of the center line of the field in each rectangular frame image.
4. The method for extracting text from a bill image according to claim 1, characterized in that After determining the sorting of each field in the corresponding column according to the ordinate of the center point of the field, it further includes: Cluster the fields in the same column based on the clustering algorithm according to the difference in the ordinate of the center point of adjacent fields in the same column, and regard the class with the largest number of fields as the second reference class; Calculate the average value of the differences in the ordinate of the center points in the second reference class to obtain a second average value, and use the second average value as the row height of the corresponding column; Judging one by one whether the difference in the ordinate of the center points of adjacent fields is greater than the line height of the corresponding column. If so, filling a second preset character between the two adjacent fields.
5. The method for extracting text from a bill image according to claim 4, wherein Performing typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and outputting each field according to the typeset page format, including: Performing typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, the sorting of each field in the corresponding column, the first preset filling character in each row, and the second preset filling character in each column, and outputting each field, each first preset filling field, and each second preset filling character according to the typeset page format.
6. The method for extracting text from a bill image according to claim 5, characterized in that, Outputting each field, each first preset filling field, and each second preset filling character according to the typeset page format, including: Outputting each field, each first preset filling field, and each second preset filling character in json format or excel file format according to the typeset page format.
7. A text extraction device for a bill image, characterized in that, Including: A field recognition module, a row determination module, a column determination module, and a typesetting module; The field recognition module is used to obtain a bill image, input the bill image into a preset text detection and recognition model, so that the text detection and recognition model recognizes each field, the abscissa of the center point of each field, and the ordinate of the center point of each field in the bill image. The row determination module is used to cluster each field according to the difference in the ordinate of the center points of each field, and then regard the fields in the same class as the fields in the same row; determining the row order of each row according to the average value of the ordinate of the center points of the fields in each row. Determining the sorting of each field in the corresponding row according to the abscissa of the center point of the field; clustering the fields in the same row based on the clustering algorithm according to the difference in the abscissa of the center points of adjacent fields in the same row, and regarding the class with the largest number of fields as the first reference class. Calculating the average value of the differences in the abscissa of the center points in the first reference class to obtain a first average value, and using the first average value as the column width of the corresponding row; judging one by one whether the difference in the abscissa of the center points of adjacent fields is greater than the column width of the corresponding row. If so, filling a first preset character between the two adjacent fields. The column determination module is used to cluster each field according to the difference in the abscissa of the center points of each field, and then regard the fields in the same class as the fields in the same column; determining the column order of each column according to the average value of the abscissa of the center points of the fields in each column. Determining the sorting of each field in the corresponding column according to the ordinate of the center point of the field; The typesetting module is used to perform typesetting according to the row order of each row, the column order of each column, the sorting of each field in the corresponding row, and the sorting of each field in the corresponding column, and output each field according to the typeset page format.
8. The text extraction device for a bill image according to claim 7, characterized in that, It also includes: Image preprocessing module, which is used to perform image preprocessing on the bill image before the field recognition module inputs the bill image into a preset text detection and recognition model; wherein, the image preprocessing includes any one or a combination of the following: removing blurred images, correcting the bill angle, and removing the bill official seal.
Citation Information
Patent Citations
Method of recognizing layout of document image
CN104966051A
Spatial position-based method for reconstructing text information in CAD electronic data
CN105808511A
Method and device for determining position of text area in image, equipment and storage medium
CN110852229A
Character processing and identifying methods, storage medium, and terminal device
CN113557520A