Image text recognition methods, devices and electronic equipment
By obtaining the matrix of the target image and the recognition template for correction and recognition, the problem of low recognition accuracy caused by inconsistent image quality of the waybill is solved, thus improving the efficiency and accuracy of waybill entry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN SANY INTELLIGENT CONTROL EQUIP
- Filing Date
- 2023-07-27
- Publication Date
- 2026-07-17
AI Technical Summary
During concrete transportation, the inconsistent quality of waybill images leads to low accuracy in OCR text recognition, affecting the efficiency of waybill entry.
By acquiring the target image, a target matrix and recognition template are obtained, image correction is performed, the region to be recognized is selected, and the OCR algorithm is applied to recognize the text content.
It improved the accuracy of text recognition and input efficiency on waybills, and reduced the error rate.
Smart Images

Figure CN117152758B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, specifically to an image text recognition method, apparatus, and electronic device. Background Technology
[0002] During concrete transportation, it is necessary to enter concrete transport manifests. Related technologies typically employ image recognition plugins, utilizing OCR text recognition technology to identify the text content on the manifests. However, due to the inconsistent quality of the manifest images taken by staff, the accuracy of text recognition is often low, affecting the efficiency of manifest entry. Summary of the Invention
[0003] To address the aforementioned technical problems, embodiments of this application provide an image text recognition method, apparatus, and electronic device, which can improve the accuracy of text content recognition on waybills and increase the efficiency of waybill entry.
[0004] Firstly, an image text recognition method is provided, including:
[0005] Acquire the target image;
[0006] Based on the target image, a target matrix and a recognition template corresponding to the target image are obtained; wherein, the target matrix represents the structure and size of the table in the preset image;
[0007] The target image is corrected based on the target image and the target matrix;
[0008] Using the recognition template, the region to be recognized is selected in the corrected target image; and
[0009] Identify the text content within the area to be identified.
[0010] According to a first aspect of this application, obtaining a target matrix and a recognition template corresponding to the target image based on the target image includes:
[0011] Obtain the category of the target image; and
[0012] Obtain the target matrix and the recognition template corresponding to the category of the target image.
[0013] According to a first aspect of this application, the correction of the target image based on the target image and the target matrix includes:
[0014] Based on the target image and the target matrix, a position matrix is obtained; wherein, the position matrix represents the measured relative positional relationship of different nodes of the table in the target image relative to a preset origin; and
[0015] The tables in the target image are corrected based on the target matrix and the position matrix.
[0016] According to a first aspect of this application, the target matrix includes a node matrix and a reference matrix; wherein, the node matrix represents the table structure in the preset image; and the reference matrix represents the original relative positional relationship of different nodes of the table in the preset image relative to the preset origin.
[0017] The step of correcting the tables in the target image based on the target matrix and the position matrix includes:
[0018] The maximum border of the table in the target image is corrected based on the reference matrix and the position matrix; and / or
[0019] The cells of the table in the target image are corrected based on the node matrix, the reference matrix, and the position matrix.
[0020] According to a first aspect of this application, correcting the cells of a table in the target image based on the node matrix, the reference matrix, and the position matrix includes:
[0021] Based on the node matrix, search for the coordinates of the four nodes corresponding to the cell on the node matrix; and
[0022] The cell is corrected based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix.
[0023] According to a first aspect of this application, the step of correcting the cell based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix includes:
[0024] Based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix, the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix are obtained.
[0025] The cell is corrected based on the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix.
[0026] According to a first aspect of this application, obtaining the position matrix based on the target image and the target matrix includes:
[0027] Based on the target image, obtain the relative coordinate values of multiple nodes in the target image relative to the preset origin; and
[0028] The position matrix is obtained based on the target matrix and the multiple relative coordinate values.
[0029] According to a first aspect of this application, the identification template includes a positioning frame and a marker frame; wherein the positioning frame is used to obtain position information of a reference area; and the marker frame is used to obtain position information of the area to be identified.
[0030] The process of selecting the region to be identified in the corrected target image by applying the recognition template includes:
[0031] Based on the recognition template, obtain the current reference character coordinates within the reference area of the positioning frame;
[0032] The deviation between the preset reference character coordinates and the current reference character coordinates is obtained based on the preset reference character coordinates and the current reference character coordinates of the positioning frame;
[0033] Adjust the position of the marker box according to the deviation amount; and
[0034] The area selected by the adjusted marker box is taken as the area to be identified.
[0035] According to a first aspect of this application, before correcting the target image based on the target image and the target matrix, the image text recognition method further includes:
[0036] The target image is rotated as a whole.
[0037] According to a first aspect of this application, the overall rotation of the target image includes:
[0038] The target image is projected and transformed within a preset angle range using a preset angle step size to obtain the projection integral corresponding to different angles.
[0039] Obtain the target angle corresponding to the maximum projection integral; and
[0040] The target image is rotated as a whole according to the target angle.
[0041] According to a first aspect of this application, the overall rotation of the target image includes:
[0042] Run the inversion detection model;
[0043] If the inversion detection model outputs a signal indicating that the target image is upside down, the target image is rotated 180 degrees.
[0044] Secondly, an image text recognition device is also provided, comprising:
[0045] The first acquisition module is configured to acquire the target image;
[0046] The matching module is configured to obtain a target matrix and a recognition template corresponding to the target image based on the target image; wherein the target matrix represents the structure and size of a table in a preset image;
[0047] The first correction module is configured to correct the target image based on the target image and the target matrix;
[0048] The first selection module is configured to apply the recognition template to select the region to be recognized in the corrected target image; and
[0049] The recognition module is configured to recognize the text content within the area to be recognized.
[0050] Thirdly, an electronic device is also provided, comprising:
[0051] processor;
[0052] And a memory for storing the processor's executable instructions;
[0053] The processor is used to execute the image text recognition method described in the above embodiments.
[0054] The image text recognition method, apparatus, and electronic device provided in this application acquire a target image, then obtain a target matrix and recognition template corresponding to the target image, then correct the target image based on the target image and target matrix, and then apply the recognition template to select the region to be recognized in the corrected target image and recognize the text content within the region. Thus, firstly, correcting the target image can adjust the position of the content (tables, text, etc.) in the target image, which is beneficial for accurate subsequent recognition of the text content in the target image, effectively improving the efficiency of entering documents such as invoices and prescriptions; secondly, applying the recognition template can more accurately determine the position of the region to be recognized in the target image, reducing the likelihood of errors in the recognition of the region's position, thereby lowering the error rate of text content recognition, improving the accuracy of text content recognition, and effectively improving the efficiency of entering documents such as invoices and prescriptions. Attached Figure Description
[0055] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0056] Figure 1 This is a flowchart illustrating an exemplary embodiment of the image text recognition method provided in this application.
[0057] Figure 2 This is a schematic diagram illustrating the process of obtaining a target matrix and a recognition template corresponding to a target image based on a target image, provided as an exemplary embodiment of this application.
[0058] Figure 3 This is a schematic diagram illustrating a process for correcting a target image based on a target image and a target matrix, provided as an exemplary embodiment of this application.
[0059] Figure 4 A schematic diagram of a table provided for an exemplary embodiment of this application.
[0060] Figure 5 This is a schematic diagram of a node matrix provided for an exemplary embodiment of this application.
[0061] Figure 6 A schematic diagram of a reference matrix provided for an exemplary embodiment of this application.
[0062] Figure 7 This is a schematic diagram illustrating the process of obtaining a position matrix based on a target image and a target matrix, provided as an exemplary embodiment of this application.
[0063] Figure 8 This is a schematic flowchart illustrating the process of correcting tables in a target image based on a target matrix and a position matrix, provided as an exemplary embodiment of this application.
[0064] Figure 9 This is a schematic diagram illustrating a process for correcting the cells of a table in a target image based on a node matrix, a reference matrix, and a position matrix, as provided in an exemplary embodiment of this application.
[0065] Figure 10 This is a flowchart illustrating a process for correcting a cell based on the coordinates of the four nodes corresponding to the cell, a reference matrix, and a position matrix, as provided in an exemplary embodiment of this application.
[0066] Figure 11 A flowchart illustrating the process of correcting cells, provided as an exemplary embodiment of this application.
[0067] Figure 12 This is a schematic diagram illustrating the process of selecting the region to be identified in a corrected target image using an application recognition template provided as an exemplary embodiment of this application.
[0068] Figure 13 A flowchart illustrating an image text recognition method provided as another exemplary embodiment of this application.
[0069] Figure 14 This is a schematic diagram illustrating the process of rotating a target image as a whole, provided as an exemplary embodiment of this application.
[0070] Figure 15 This is a schematic diagram illustrating the process of rotating a target image as a whole, provided as another exemplary embodiment of this application.
[0071] Figure 16 This is a structural block diagram of an image text recognition device provided for an exemplary embodiment of this application.
[0072] Figure 17 This is a structural block diagram of an image text recognition device provided for an exemplary embodiment of this application.
[0073] Figure 18 A structural block diagram of an electronic device provided for an exemplary embodiment of this application. Detailed Implementation
[0074] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0075] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of the image text recognition method provided in this application. Figure 1 As shown, the image text recognition method provided in this application embodiment may include:
[0076] S210: Acquire the target image.
[0077] In one embodiment, the target image may include invoices, prescriptions, expense receipts, etc.
[0078] In one embodiment, the content in the target image may include tables, text content, etc.
[0079] S220: Based on the target image, obtain the target matrix and recognition template corresponding to the target image.
[0080] Specifically, the target matrix can represent the structure and size of the table in the preset image. It should be noted that the preset image and the target image can be images of the same category; the preset image is the standard image corresponding to the target image under standard shooting conditions. In practical applications, the preset image, along with the bound target matrix and recognition template, are pre-stored in the system. Based on the target image, the corresponding preset image, as well as the target matrix and recognition template bound to the preset image, can be determined. The structure of the table can be understood as the distribution of multiple cells in the table, and the size of the table can be understood as the relative positional relationship of multiple nodes in the table relative to the preset origin. It should be understood that the multiple nodes of the table are the intersection points of multiple horizontal lines and multiple vertical lines in the table.
[0081] Specifically, the recognition module can be used to mark the area to be recognized where the text content in the target image is located, which is beneficial for the accurate subsequent input of the corresponding text content.
[0082] S230: Correct the target image based on the target image and the target matrix.
[0083] Specifically, during step S210, the shooting angle may cause the content (tables, text, etc.) in the target image to shift or tilt. Step S230 corrects the target image based on the target matrix, adjusting the position of the content (tables, text, etc.) to facilitate accurate identification of text content and effectively improve the efficiency of data entry for documents such as invoices and prescriptions. The specific correction process will be described in detail later.
[0084] S240: Apply the recognition template to select the region to be recognized in the corrected target image.
[0085] S250: Recognizes text content within the area to be recognized.
[0086] Specifically, the recognition template can mark the area where the text content in the target image is located, i.e., the area to be recognized. After the location of the area to be recognized is determined, step S250 is executed, and the OCR recognition algorithm is called to recognize the text content in the area marked by the recognition template, and the corresponding text content can be entered.
[0087] It should be understood that by executing step S240, a recognition template can be applied to more accurately determine the area where the text content in the target image is located. This reduces the likelihood of errors in the location recognition of the area to be recognized, thereby lowering the error rate of text content recognition, improving the accuracy of text content recognition, and effectively improving the entry efficiency of the aforementioned documents such as invoices and prescriptions.
[0088] The image text recognition method provided in this application involves acquiring a target image, obtaining a target matrix and a recognition template corresponding to the target image, correcting the target image based on the target image and the target matrix, and then applying the recognition template to select the region to be recognized in the corrected target image and recognize the text content within the region. This achieves two main benefits: First, correcting the target image adjusts the position of content (tables, text, etc.) within the target image, facilitating accurate subsequent recognition of the text content and effectively improving the efficiency of entering documents such as invoices and prescriptions. Second, applying the recognition template more accurately determines the region where the text content is located in the target image, reducing the likelihood of location errors and lowering the error rate of text content recognition, thus improving the accuracy of text content recognition and effectively increasing the efficiency of entering documents such as invoices and prescriptions.
[0089] Figure 2 This is a schematic diagram illustrating a process for obtaining a target matrix and a recognition template corresponding to a target image based on a target image, as provided in an exemplary embodiment of this application. Figure 2 As shown, step S220 may include:
[0090] S221: Obtain the category of the target image.
[0091] S222: Obtain the target matrix and recognition template corresponding to the category of the target image.
[0092] It should be noted that the table structure and size in the target image will be different depending on the category of the target image, and the corresponding preset image category will also be different. The target matrix and recognition module bound to the preset image will also be different. By executing steps S221 and S222, the preset image of the corresponding category can be obtained according to the target image of different categories, thereby obtaining the corresponding target matrix and recognition module, which facilitates the subsequent targeted correction of the target image of the corresponding category.
[0093] In one embodiment, a deep learning image classification method can be used to classify preset images. Specifically, the deep learning network can be a ResNet18 residual convolutional neural network. Before executing step S221, 60 preset images of each type can be prepared as a dataset, and these 60 preset images of each type can be divided according to different orientations of the image layout. For example, they can be divided into 15 images each for orientations of 0 degrees, 90 degrees, 180 degrees, and 270 degrees. Then, classification training is performed on all types of preset images. After training, the category label of each preset image is bound to the target matrix and recognition template defined for each preset image. In this way, during the execution of steps S221 and S222, by identifying the category of the target image, preset images of the same category as the target image can be retrieved, thereby obtaining the corresponding target matrix and recognition template.
[0094] In one embodiment, if all target images are of the same type, it is not necessary to determine the category of the target image. Instead, the corresponding preset image, as well as the target matrix and recognition template bound to the preset image, can be obtained directly based on the current target image.
[0095] Figure 3 This is a schematic diagram illustrating a process for correcting a target image based on a target image and a target matrix, provided as an exemplary embodiment of this application. Figure 3 As shown, step S230 includes:
[0096] S231: Obtain the position matrix based on the target image and the target matrix.
[0097] S232: Correct the tables in the target image based on the target matrix and the position matrix.
[0098] Specifically, the target matrix may include a node matrix, which can represent the table structure in the preset image. That is, the node matrix can mathematically express different types of tables, and the distribution position of different nodes of the table in the preset image can be determined through the node matrix.
[0099] For example, Figure 4 A schematic diagram of a table provided for an exemplary embodiment of this application. Figure 4 The table in the image has 10 rows and 9 columns. Correspondingly, the node matrix is 10x9 in size. The rows and columns of the node matrix correspond to the rows and columns of the table in the target image. The values of the matrix elements in the node matrix can correspond to the number of line segments on the table nodes, or they can be 0 or 1 to indicate whether a node exists at a corresponding position in the table. This article uses the example of the matrix element values corresponding to the number of line segments on the table nodes to introduce this node matrix. Figure 4The table shown has a corresponding node matrix A[10, 9]. The number of line segments on the first node in the top left corner of the table is 2, so the corresponding node matrix A[0, 0] = 2. If there is no node at the table position corresponding to a matrix element in the node matrix, then the value of that matrix element is 0. Figure 4 In the given information, A[0, 2] = 0, and so on, we can... Figure 4 The tables in the table are represented by a node matrix A. Figure 5 This is a schematic diagram of a node matrix provided for an exemplary embodiment of this application. Figure 5 As shown, Figure 5 The node matrix shown can be understood as Figure 4 The table in the table is a matrix obtained through the aforementioned transformation rules.
[0100] Specifically, the target matrix may also include a reference matrix, which represents the original relative positional relationship of different nodes in the table in the preset image relative to the preset origin. The size of the reference matrix is the same as that of the aforementioned node matrix, that is, the rows and columns of the reference matrix correspond to the rows and columns of the table in the preset image. Thus, the reference matrix can be defined as B[10, 9], and the first node in the upper left corner of the table is defined as the preset origin position, i.e., B[0, 0] = (0, 0). The position of the matrix element 0 in the node matrix A is defined as (-1, -1) for the corresponding position element in the reference matrix B, for example, B[0, 2] = (-1, -1); while the position of the matrix element that is not 0 in the node matrix A is the distance value of the node relative to the preset origin in the row and column directions, for example, B[0, 1] = (0, 125). It should be noted that the distance values of the node relative to the preset origin in the row and column directions are the distance values measured under normal conditions (without deviation or offset due to shooting), which are the original standard values. Figure 6 This is a schematic diagram of a reference matrix provided for an exemplary embodiment of this application. Figure 6 As shown, Figure 6 The reference matrix shown can be understood as Figure 5 The table in the table is a matrix obtained through the aforementioned transformation rules.
[0101] Specifically, the position matrix can represent the actual relative positional relationship of different nodes in the table in the target image relative to the preset origin. The position matrix has the same size as the aforementioned node matrix, that is, the rows and columns of the position matrix correspond to the rows and columns of the table in the target image. Thus, the position matrix can be defined as C[10, 9]. Based on the distribution of the node matrix, for the position where the matrix element in the node matrix is 0, the value of the corresponding position element in the position matrix C is defined as (-1, -1), for example, C[0, 2] = (-1, -1). This can improve the generation efficiency of the position matrix. For the position where the matrix element in the node matrix A is not 0, the position matrix C can be obtained by taking the actual measured distance values of each node relative to the preset origin in the row and column directions in the current target image (the target image obtained by executing step S210 after actual shooting, where the table content is deviated and tilted). These are the actual measured values, and so on, to obtain the position matrix C.
[0102] In one embodiment, the position matrix is the same size as the aforementioned reference matrix, that is, the rows and columns of the position matrix correspond to the rows and columns of the table in the target image. Alternatively, the corresponding position matrix can be generated using the distribution content of the reference matrix.
[0103] It should be understood that by executing step S232, the node matrix can be used to search for the position of nodes in the target image. By comparing the reference matrix with the position matrix, the deviation value of the current target image relative to the normal situation (without deviation due to shooting) can be obtained. In this way, the target matrix (including the node matrix and the reference matrix) can be used to more accurately correct the table in the target image, which is beneficial to improving the recognition accuracy of text content.
[0104] Figure 7 This is a schematic diagram illustrating a process for obtaining a position matrix based on a target image and a target matrix, provided as an exemplary embodiment of this application. Figure 7 As shown, step S231 includes:
[0105] S2311: Based on the target image, obtain the relative coordinates of multiple nodes in the target image with respect to the preset origin.
[0106] S2312: Obtain the position matrix based on the target matrix and multiple relative coordinate values.
[0107] Specifically, after obtaining the target image, it can first be converted into a grayscale image. Then, the Gaussian filter operator from the OpenCV library is used to perform Gaussian filtering on the image using a Gaussian filter convolution kernel. The filtered image is then processed using OpenCV's Canny operator to obtain the binary image of the target image's contour region. Morphological opening operations and contour shape selection methods are then used to process the binary image, filtering out the regions containing OCR-recognized text, thus obtaining the binary image of the table in the target image. Finally, the image morphological opening operation is used again to obtain the binary images of the horizontal and vertical lines of the table.
[0108] It should be noted that when acquiring the binary image of the horizontal lines of the table, the width of the convolution kernel for the morphological opening operation is set to twice the width of the table line in pixels, and the height is set to 1; when acquiring the binary image of the vertical lines of the table, the width of the convolution kernel for the morphological opening operation is set to 1, and the height is set to twice the width of the line in pixels. Using the acquired binary images of the horizontal and vertical lines of the table, the `bitwise` function of OpenCV is used to obtain the binary image of the intersection points of the horizontal and vertical lines. After using the `findContours` function of OpenCV to perform contour finding on the intersection point binary image, the midpoint of each contour is obtained, which is the intersection point of the horizontal and vertical lines of the table, i.e., the node of the table in the current target image. The coordinate values of all nodes are then converted relative to a preset origin (…). Figure 4 The relative coordinate values of the first node in the top left corner of the table are obtained, and based on the distribution of the matrix elements of the aforementioned node matrix, all relative coordinate values are summarized into a position matrix of the table. According to the distribution of the target matrix (the position of the element with a value of (-1, -1) in the reference matrix or the position of the element with a value of 0 in the node matrix), the corresponding position in the position matrix is assigned the value (-1, -1). In this way, the efficiency of assigning values to the position matrix can be improved.
[0109] Figure 8 This is a schematic flowchart illustrating the process of correcting tables in a target image based on a target matrix and a position matrix, provided as an exemplary embodiment of this application. Figure 8 As shown, step S232 includes:
[0110] S2321: Correct the maximum border of the table in the target image based on the reference matrix and the position matrix.
[0111] Specifically, based on the aforementioned position matrix, the node coordinate set V1 (C[0,0], C[0,8], C[9,8], C[9,0]) of the maximum border of the table in the current target image (the target image obtained by executing step S210) can be obtained. Based on the aforementioned reference matrix, the node coordinate set V2 (B[0,0], B[0,8], B[9,8], B[9,0]) of the maximum border of the table in the target image under normal conditions (the target image without deviation due to shooting) can be obtained. Then, based on V1 and V2, the planar perspective transformation matrix M1 from V1 to V2 can be calculated. Based on M1, global perspective correction is performed on the current target image. In this way, the position of the maximum border of the table in the target image can be corrected.
[0112] S2322: Correct the cells of the table in the target image based on the node matrix, reference matrix, and position matrix.
[0113] Specifically, the table in the target image comprises multiple cells. Step S2322 allows for individual correction of the position of each cell, thereby effectively correcting the entire table in the target image. The specific correction process for each cell will be detailed later.
[0114] In one embodiment, step S2322 can be performed by correcting only one or two cells of the table in the target image, or by correcting multiple cells of the table sequentially.
[0115] It should be noted that in practical applications, steps S2321 and S2322 can both be executed, or only one of them can be executed.
[0116] Figure 9 This is a schematic diagram illustrating a process for correcting the cells of a table in a target image based on a node matrix, a reference matrix, and a position matrix, as provided in an exemplary embodiment of this application. Figure 9 As shown, step S2322 may include:
[0117] S23221: Based on the node matrix, search for the coordinates of the four nodes corresponding to the cell in the node matrix.
[0118] Specifically, based on the node matrix, the positions of nodes in the table can be searched along the row or column direction. For example, using... Figure 4Taking the first cell in the top left corner as an example, using the rows of the node matrix as a reference, starting from the preset origin of the node matrix, search the matrix elements in different columns of the first row. If the search result indicates that there is no node in the current column of the first row, continue searching the matrix elements in the next column along the direction of the first row. If the search result indicates that there is a node in the current column of the first row, then obtain the position coordinates of the node based on the matrix elements of the node matrix. Then search the matrix elements in the next row of the same column as the node. If the search result indicates that there is no node in the current column of the second row, continue searching the matrix elements in the next row along the direction of the current column. If the search result indicates that there is a node in the current column of the second row, then the position coordinates of the node can be obtained by clicking the matrix elements. In this way, relying on the node matrix, the nodes of each cell can be traversed to obtain the coordinates of the four nodes corresponding to each cell.
[0119] S23222: Correct the cell based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix.
[0120] Specifically, Figure 10 This is a schematic diagram illustrating a process for correcting a cell based on the coordinates of the four nodes corresponding to the cell, a reference matrix, and a position matrix, as provided in an exemplary embodiment of this application. Figure 10 As shown, step S23222 includes:
[0121] S232221: Based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix, obtain the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix.
[0122] S232222: Correct the cell based on the coordinate set of the four nodes corresponding to the cell in the position matrix and the coordinate set in the reference matrix.
[0123] Specifically, with Figure 4 Taking the first cell in the top left corner as an example, after obtaining the coordinates of multiple nodes in the cell through the node matrix, the coordinate set V3 (C[0,0], C[0,1], C[1,1], C[1,0]) of multiple nodes on the position matrix C and the coordinate set V4 (B[0,0], B[0,1], B[1,1], B[1,0]) on the reference matrix B can be found based on the coordinates of multiple nodes. Then, based on V3 and V4, the planar perspective transformation matrix M2 from V3 to V4 can be calculated. Based on M2, the current cell is corrected for perspective. In this way, by executing steps S23221 and S23222 multiple times, the position of each cell in the table in the target image can be corrected.
[0124] Specifically, Figure 11This is a flowchart illustrating a cell correction procedure provided for an exemplary embodiment of this application. The process of repeatedly executing steps S23221 and S23222 can be implemented by setting a corresponding algorithm program, see reference... Figure 11 The program diagram shown illustrates the program execution process as follows:
[0125] S1: On node matrix A, using the rows of the node matrix as the reference, search according to the columns of the matrix. Define the coordinates of the first node in the upper left corner of the cell as A[m, n]. When starting the search, set m = 0, n = 0, and define a = 0.
[0126] S2: Search by column, determine if the value of the next column of the current row of node matrix A is 0, i.e., a = a + 1, and whether A[m, n + a] is greater than 0. If it is not greater than 0, continue to execute S2; otherwise, execute S3.
[0127] S3: Determine whether the point in the next row corresponding to point A[m, n+a] is greater than 0, that is, whether A[m+1, n+a] is greater than 0. If it is not greater than 0, execute S2; otherwise, execute S4.
[0128] S4: Let the coordinates of the second point of the cell be A[m, n+a], the coordinates of the third point be A[m+1, n+a], and the coordinates of the fourth point be A[m+1, n]. Based on the relationship between the node matrix A, the reference matrix B, and the position matrix C, the coordinate set of the cell on the position matrix C is V3[C[m, n], C[m, n+a], C[m+1, n+a], C[m+1, n]], and the coordinate set on the reference matrix B is V4[B[m, n], B[m, n+a], B[m+1, n+a], B[m+1, n]]. Calculate the planar perspective transformation matrix M2 from V3 to V4, and then use the perspective transformation matrix M2 to perform perspective correction on the image of the corresponding cell area to obtain the corrected image of the cell area.
[0129] S5: Let n = n + a. Check if the coordinate point A[m, n+1] exists in the node matrix. If it exists, record A[m, n] as the first coordinate point of the cell and execute S2. Otherwise, execute S6.
[0130] S6: Determine if coordinate point A[m+2, 0] exists in the node matrix. If it exists, m = m+1, n = 0, a = 0, and record A[m, n] as the first coordinate point of the cell. Then execute S2. Otherwise, the cell correction ends.
[0131] Figure 12 This is a schematic diagram illustrating the process of selecting a region to be identified in a corrected target image using an application recognition template, as provided in an exemplary embodiment of this application. Figure 12 As shown, step S240 includes:
[0132] S241: Based on the recognition template, obtain the current reference character coordinates within the reference area of the positioning box.
[0133] S242: Based on the preset reference character coordinates and the current reference character coordinates of the positioning frame, obtain the deviation between the preset reference character coordinates and the current reference character coordinates.
[0134] Specifically, the recognition template may include a positioning box, which can be used to obtain the position information of the reference area. When applying the recognition template, for different target images, the positioning box can obtain the position information of the reference area corresponding to the different target images, thereby playing a positioning role and confirming whether the target image is within the area of OCR text recognition.
[0135] It should be noted that during the creation of the recognition template, the created positioning box is positioned within the reference area of the target image under normal conditions. The preset reference character coordinates within this reference area are obtained and pre-stored in the system as a reference standard. After executing step S241, the recognition module can obtain the current reference character coordinates within the reference area of the positioning box. Then, by executing step S242, the deviation between the preset reference character coordinates and the current reference character coordinates can be obtained.
[0136] In one embodiment, during the creation of the recognition template, a region in different target images with less variation in text content is selected as a reference region for positioning within the positioning box.
[0137] In one embodiment, in order to enable the positioning frame to recognize reference characters within a larger range, the width of the positioning frame is typically set to 1.1 to 1.2 times the width of the reference area, and the height of the positioning frame is set to 3 times the height of the reference area during the creation of the recognition template.
[0138] S243: Adjust the position of the marker box according to the deviation amount.
[0139] S244: Use the area selected by the adjusted marker box as the area to be identified.
[0140] Specifically, the recognition template may include a bounding box, which can be used to obtain the position information of the area to be recognized. For the same recognition template, the relative position between the positioning box and the bounding box will not change. Therefore, after obtaining the deviation between the preset reference character coordinates and the current reference character coordinates, step S243 can be executed to adjust the position of the bounding box according to the deviation. This can more accurately adjust the bounding box to the area where the text content is located, i.e. the area to be recognized, so that the OCR algorithm can more accurately recognize the text content of the target image within the bounding box.
[0141] In one embodiment, in order to enable the marker box to better select the text content within the area to be recognized, the width and height of the marker box are usually set to 1.1 to 1.2 times the width and height of the area to be recognized during the creation of the recognition template.
[0142] Figure 13 This is a flowchart illustrating an image text recognition method provided as another exemplary embodiment of this application. Figure 13 As shown, before step S230, the image text recognition method further includes:
[0143] S260: Rotate the target image as a whole.
[0144] Specifically, since the target image may be tilted or upside down when the template image is captured, step S260 is performed to rotate the target image as a whole (including tilt correction and upside down correction) before recognizing the text content of the target image. This can bring the target image to a roughly normal state. Then step S230 is performed to correct the content of the target image. This makes it easier to correct the target image in the future and is more conducive to improving the accuracy of recognizing the text content in the target image.
[0145] Figure 14 This is a schematic diagram illustrating the process of rotating a target image as a whole, provided as an exemplary embodiment of this application.
[0146] like Figure 14 As shown, step S260 includes:
[0147] S261: Perform projection transformation on the target image within a preset angle range with a preset angle step size to obtain the projection integral corresponding to different angles.
[0148] S262: Obtain the target angle corresponding to the maximum projection integral.
[0149] S263: Rotate the target image as a whole according to the target angle.
[0150] It should be noted that the preset angle step size can be set according to actual conditions, and the embodiments of this application do not specifically limit the preset angle step size. In one embodiment, the preset angle step size can be selected as 1 degree, 2 degrees, 3 degrees, etc.
[0151] It should be noted that the preset angle range can be set according to the actual situation, and the embodiments of this application do not specifically limit the preset angle range. In one embodiment, the preset angle range can be selected as 0-179 degrees, 0-169 degrees, 10-179 degrees, etc.
[0152] Specifically, the Radon transform algorithm can be used to perform a projection transformation on the target image within a preset angle range (e.g., from 0 degrees to 179 degrees). The projection integral is calculated with a preset angle step size (e.g., 1 degree each time) to obtain 180 projection integral values. The angle corresponding to the largest value among the 180 projection integral values is the rotation angle of the current image, defined as d. After rotating the image by an angle of -d, the target image with the text content remaining horizontal can be obtained.
[0153] Figure 15 This is a schematic diagram illustrating the process of rotating a target image as a whole, provided as another exemplary embodiment of this application. Figure 15 As shown, step S260 further includes:
[0154] S264: Run the inversion detection model.
[0155] S265: If the inversion detection model outputs a signal representing the upside down of the target image, rotate the target image by 180 degrees.
[0156] In one embodiment, steps S264 and S265 can be executed after step S263. It should be noted that after step S263, although the text content in the target image remains horizontal overall, it may be upside down. Therefore, an inversion detection model can be run to confirm whether the current target image is upside down. If the inversion detection model outputs a signal indicating that the target image is upside down, the target image can be rotated 180 degrees to swap the positions of the upper and lower halves. This is more conducive to the subsequent accurate recognition of the text content.
[0157] In one embodiment, an inversion detection model can be obtained through deep learning training. For example, 1,000 images of different targets can be collected and made into 500 images at 0 degrees and 500 images at 180 degrees, resulting in a 0-degree image dataset and a 180-degree image dataset. The ResNet18 neural network, cross-entropy loss function, and stochastic gradient descent optimizer are used to train the two classifications using deep learning to obtain a deep learning binary classification detection model, namely the aforementioned inversion detection model.
[0158] In one embodiment, steps S261, S262 and S263 can be skipped, and steps S264 and S265 can be executed directly.
[0159] Figure 16 This is a structural block diagram of an image text recognition device provided for an exemplary embodiment of this application. (See diagram below.) Figure 16As shown, the image text recognition device 400 provided in this application embodiment may include: a first acquisition module 410, configured to acquire a target image; a matching module 420, configured to obtain a target matrix and a recognition template corresponding to the target image based on the target image; wherein the target matrix represents the structure and size of a table in a preset image; a first correction module 430, configured to correct the target image based on the target image and the target matrix; a first selection module 440, configured to apply the recognition template to select a region to be recognized in the corrected target image; and a recognition module 450, configured to recognize the text content within the region to be recognized.
[0160] The image text recognition device provided in this application acquires a target image, then obtains a target matrix and a recognition template corresponding to the target image, corrects the target image based on the target matrix, and then applies the recognition template to select the region to be recognized in the corrected target image and recognize the text within the region. This achieves two main benefits: First, correcting the target image adjusts the position of content (tables, text, etc.) within the target image, facilitating accurate subsequent recognition of text content and effectively improving the efficiency of entering documents such as invoices and prescriptions. Second, applying the recognition template more accurately determines the position of the region to be recognized in the target image, reducing the likelihood of errors in recognizing the region's position and thus lowering the error rate of text content recognition, thereby improving the accuracy of text content recognition and effectively increasing the efficiency of entering documents such as invoices and prescriptions.
[0161] Figure 17 This is a structural block diagram of an image text recognition device provided for an exemplary embodiment of this application. (See diagram below.) Figure 17 As shown, in one embodiment, the matching module 420 includes a category acquisition module 421 configured to acquire the category of the target image; and a category matching module 422 configured to acquire a target matrix and a recognition template corresponding to the category of the target image.
[0162] like Figure 17 As shown, in one embodiment, the first correction module 430 includes a first matrix generation module 431, configured to obtain a position matrix based on the target image and the target matrix; wherein the position matrix represents the actual relative positional relationship of different nodes of the table in the target image relative to a preset origin; and a second correction module 432, configured to correct the table in the target image based on the target matrix and the position matrix.
[0163] like Figure 17As shown, in one embodiment, the second correction module 432 includes a third correction module 4321 configured to correct the maximum border of the table in the target image according to the reference matrix and the position matrix; and a fourth correction module 4322 configured to correct the cells of the table in the target image according to the node matrix, the reference matrix and the position matrix.
[0164] like Figure 17 As shown, in one embodiment, the fourth correction module 4322 includes a search module 43221, configured to search for the coordinates of the four nodes corresponding to the cell on the node matrix according to the node matrix; and a fifth correction module 43222, configured to correct the cell according to the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix.
[0165] like Figure 17 As shown, in one embodiment, the fifth correction module 43222 includes a coordinate set output module 432221, configured to obtain the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix; the sixth correction module 432222 is configured to correct the cell based on the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix.
[0166] like Figure 17 As shown, in one embodiment, the first matrix generation module 431 includes a second acquisition module 4311, configured to acquire relative coordinate values of multiple nodes in the target image relative to a preset origin based on the target image; and a second matrix generation module 4312, configured to obtain a position matrix based on the target matrix and the multiple relative coordinate values.
[0167] like Figure 17 As shown, in one embodiment, the recognition template includes a positioning box and a marker box; wherein, the positioning box is used to obtain the position information of a reference area; the marker box is used to obtain the position information of the area to be recognized; the first selection module 440 includes a third acquisition module 441, configured to obtain the current reference character coordinates within the reference area of the positioning box according to the recognition template; an deviation calculation module 442, configured to obtain the deviation between the preset reference character coordinates and the current reference character coordinates according to the preset reference character coordinates of the positioning box and the current reference character coordinates; an adjustment module 443, configured to adjust the position of the marker box according to the deviation; and a second selection module 444, configured to use the area selected by the adjusted marker box as the area to be recognized.
[0168] like Figure 17 As shown, in one embodiment, the image text recognition device 400 includes a first rotation module 460 configured to rotate the target image as a whole.
[0169] like Figure 17 As shown, in one embodiment, the first rotation module 460 includes a projection transformation module 461, configured to perform projection transformation on the target image within a preset angle step size to obtain projection integrals corresponding to different angles; a fourth acquisition module 462, configured to acquire the target angle corresponding to the maximum projection integral; and a second rotation module 463, configured to rotate the target image as a whole according to the target angle.
[0170] like Figure 17 As shown, in one embodiment, the first rotation module 460 includes an inversion detection module 464 configured to run an inversion detection model; and a third rotation module 465 configured to rotate the target image by 180 degrees if the inversion detection model outputs a signal representing that the target image is upside down.
[0171] Figure 18 This is a structural block diagram of an electronic device provided as an exemplary embodiment of this application. (See diagram below.) Figure 18 As shown, the electronic device 600 provided in this application embodiment can be any one or both of the first device and the second device, or a standalone device independent of them. The standalone device can communicate with the first device and the second device to receive the collected input signals from them.
[0172] like Figure 18 As shown, the electronic device 600 includes one or more processors 610 and a memory 620. The memory 620 is used to store executable instructions of the processor 610, which is used to perform the image text recognition method as described in the previous embodiment.
[0173] The processor 610 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 600 to perform desired functions.
[0174] The memory 620 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 610 may execute the program instructions to implement the control methods and / or other desired functions of the various embodiments of this application described above. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0175] In one example, the electronic device 600 may also include an input device 630 and an output device 640, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0176] When the controller is a standalone device, the input device 630 can be a communication network connector for receiving the acquired input signals from the first device and the second device.
[0177] In addition, the input device 630 may also include, for example, a keyboard, a mouse, etc.
[0178] The output device 640 can output various information to the outside, including determined distance information, direction information, etc. The output device 640 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0179] Of course, for the sake of simplicity, Figure 18 Only some of the components of the electronic device 600 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 600 may include any other suitable components depending on the specific application.
[0180] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0181] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0182] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0183] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0184] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0185] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0186] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. An image text recognition method, characterized in that, include: Acquire the target image; Based on the target image, a target matrix and a recognition template corresponding to the target image are obtained; wherein, the target matrix includes a node matrix and a reference matrix; wherein, the node matrix represents the table structure in a preset image, the preset image and the target image are images of the same category, and the preset image is a standard image corresponding to the target image under standard shooting conditions; the reference matrix represents the original relative position relationship of different nodes of the table in the preset image relative to a preset origin; The target image is corrected based on the target image and the target matrix; Using the recognition template, the region to be recognized is selected in the corrected target image; and Identify the text content within the area to be identified; The step of correcting the target image based on the target image and the target matrix includes: Based on the target image and the target matrix, a position matrix is obtained; wherein, the position matrix represents the actual relative positional relationship of different nodes of the table in the target image relative to a preset origin; and The tables in the target image are corrected based on the target matrix and the position matrix; The step of correcting the tables in the target image based on the target matrix and the position matrix includes: The maximum border of the table in the target image is corrected based on the reference matrix and the position matrix; The cells of the table in the target image are corrected based on the node matrix, the reference matrix, and the position matrix.
2. The image text recognition method according to claim 1, characterized in that, The step of obtaining the target matrix and recognition template corresponding to the target image based on the target image includes: Obtain the category of the target image; and Obtain the target matrix and the recognition template corresponding to the category of the target image.
3. The image text recognition method according to claim 1, characterized in that, The step of correcting the cells of the table in the target image based on the node matrix, the reference matrix, and the position matrix includes: Based on the node matrix, search for the coordinates of the four nodes corresponding to the cell on the node matrix; and The cell is corrected based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix.
4. The image text recognition method according to claim 3, characterized in that, The step of correcting the cell based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix includes: Based on the coordinates of the four nodes corresponding to the cell, the reference matrix, and the position matrix, the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix are obtained. The cell is corrected based on the coordinate set of the four nodes corresponding to the cell on the position matrix and the coordinate set on the reference matrix.
5. The image text recognition method according to claim 1, characterized in that, The step of obtaining the position matrix based on the target image and the target matrix includes: Based on the target image, obtain the relative coordinate values of multiple nodes in the target image relative to the preset origin; and obtain the position matrix based on the target matrix and the multiple relative coordinate values.
6. The image text recognition method according to claim 1, characterized in that, The recognition template includes a positioning box and a marker box; wherein, the positioning box is used to obtain the position information of the reference area; and the marker box is used to obtain the position information of the area to be recognized. The process of selecting the region to be identified in the corrected target image by applying the recognition template includes: Based on the recognition template, obtain the current reference character coordinates within the reference area of the positioning frame; The deviation between the preset reference character coordinates and the current reference character coordinates is obtained based on the preset reference character coordinates and the current reference character coordinates of the positioning frame. Adjust the position of the marker box according to the deviation amount; and The area selected by the adjusted marker box is taken as the area to be identified.
7. The image text recognition method according to claim 1, characterized in that, Before correcting the target image based on the target image and the target matrix, the image text recognition method further includes: The target image is rotated as a whole.
8. An image text recognition device, applied to the image text recognition method of claim 1, characterized in that, include: The first acquisition module is configured to acquire the target image; The matching module is configured to obtain a target matrix and a recognition template corresponding to the target image based on the target image; wherein the target matrix represents the structure and size of a table in a preset image; A first correction module is configured to correct the target image based on the target image and the target matrix; a first selection module is configured to apply the recognition template to select a region to be recognized in the corrected target image; and The recognition module is configured to recognize the text content within the area to be recognized.
9. An electronic device, characterized in that, include: processor; And a memory for storing the processor's executable instructions; The processor is used to execute the image text recognition method as described in any one of claims 1 to 7.