Business flow identification method and device based on image identification
By obtaining the combination of text line position and table detection model, the results of peer text and column line division are constructed, and the problem of extracting the table structure of wireless bank statements is solved, and fast and accurate business statement recognition is achieved.
Patent Information
- Application Number
- CN202510598133.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
AI Technical Summary
It is difficult for the prior art to quickly and accurately extract structural information in wireless bank statement tables, especially when the text dense adhesions of adjacent columns and the single item information has multiple rows up and down.
By obtaining the position of text rows and cropping the image fragments of text rows, using the recognition model to obtain the recognition results and single-word positions of text rows, combining the position annotation of the table detection model, constructing peer text and generating column line division results, calculating the numerical relationship between the two columns of text, positioning the balance column, and generating business flow results based on the virtual row and column line coordinates.
It realizes fast and efficient structural extraction of wireless flow tables, improves the accuracy and generalization of wireless flow tables, and can generate business flow results stably.
Smart Images

Figure CN120526435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a business flow recognition method and device based on image recognition. Background Art
[0002] Counting business transactions is one of the important tasks in the financial field. As imaged digital documents gradually replace paper documents as the carrier of business information, how to quickly and accurately restore the structure of bank transactions has become an important research direction.
[0003] Taking banking as an example, when a customer applies for a loan, the bank's front desk staff will scan the information into an image and input it into the system. Since the bank transaction formats of different banking businesses vary greatly, for example, some tables have table lines, while some tables have no table lines at all, for tables with table lines, the existing related technologies can stably extract their structure by constructing the structure using table lines. However, for bank transaction forms without table lines (hereinafter referred to as wireless bank transaction forms), when the text in adjacent columns is densely adhered and a single item of information exists in multiple lines above and below, how to stably segment the adhered text and merge the text lines belonging to the same information item to achieve correct structure extraction is still a very challenging task. Therefore, there is an urgent need for a method for business transaction recognition for wireless forms. Summary of the Invention
[0004] The present invention provides a business flow recognition method and device based on image recognition, which is used to solve the defect in the prior art that it is difficult to achieve correct structure extraction for wireless flow tables, and realize fast and efficient structure extraction of wireless flow tables and generate flow information.
[0005] The present invention provides a business flow recognition method based on image recognition, comprising:
[0006] Obtaining the text line position of the original image and cropping it to obtain a text line image segment, inputting the text line image segment into a recognition model, and obtaining the recognition result, single word content, and single word position of the text line output by the recognition model;
[0007] Inputting the original image into a table detection model to obtain a position mark of the table area output by the table detection model;
[0008] Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line;
[0009] Based on the individual word content and the individual word position, generating the column line division result of the text in the same row according to a preset order;
[0010] Based on the text in the same row and the column line division result, a numerical relationship between two columns of text in the same row is calculated to locate a balance column, the text in the same row corresponding to the balance column is defined as a main row, and a row line division result between adjacent main rows is generated based on the main row;
[0011] Based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells. The text corresponding to the main row is filled into the corresponding virtual cells to generate a business flow result.
[0012] According to a business flow recognition method based on image recognition provided by the present invention, based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line, including:
[0013] sorting the text lines in a preset order according to the position coordinates of the text lines;
[0014] Draw a foreground image of the text line, wherein the size of the foreground image is the same as that of the original image, the value of the text line area in the foreground image is a unique ID corresponding to each text line, and the value of other areas is 0;
[0015] According to the position mark and the sorting order of the text lines, loopingly find the nearest text line belonging to the same line on the right side of each text line and connecting them, and judging whether the connection meets the preset conditions based on the foreground image;
[0016] The text lines connected in pairs that meet the preset conditions are concatenated into the same line of text.
[0017] According to a business flow recognition method based on image recognition provided by the present invention, based on the position mark and the sort order of the text lines, the nearest text line belonging to the same line on the right side of each text line is cyclically found and connected, and whether the connection meets the preset conditions is determined based on the foreground image, including:
[0018] Connect the center points of the left side and the right side of the position mark of text line text_i, and extend the extension line to text line text_j. When the extension line passes through the edge of the position mark of text line text_j that is adjacent to text line text_i, and the contact point between the extension line and the adjacent edge divides the adjacent edge into two segments, and the ratio of the two segments is greater than a preset threshold, it is determined that text line text_i and text line text_j meet the first connection condition.
[0019] Connecting a center point of the position mark of text line text_i with a center point of the position mark of text line text_j, and determining that text line text_i and text line text_j meet a second connection condition when the coordinates of the connecting line always correspond to an area with a value of 0 in the foreground image;
[0020] In the case that the text line text_i and the text line text_j satisfy both the first connection condition and the second connection condition, it is determined that the text line text_i and the text line text_j satisfy the connection condition.
[0021] According to a business flow recognition method based on image recognition provided by the present invention, based on the content and position of the single word, a column and line division result of the same line text is generated in a preset order, including:
[0022] Based on the word content and the word position, generating candidate segmentation points for the text line, the candidate segmentation points including a head segmentation point and a tail segmentation point, the first word of the text line corresponding to the head segmentation point, and the last word corresponding to the tail segmentation point;
[0023] According to the order of the text lines in the same row from top to bottom, searching downward for similar cut points for the candidate cut points of the text lines and performing virtual connections;
[0024] Aggregating the candidate segmentation points with virtual connections into a virtual segmentation line; the virtual segmentation line includes a head segmentation line formed by a line connecting the head segmentation points and a tail segmentation line formed by a line connecting the tail segmentation points;
[0025] Performing credibility calculation on the virtual dividing lines, and removing the virtual dividing lines that do not meet preset conditions;
[0026] Sort the remaining virtual dividing lines in order from left to right;
[0027] Among the sorted virtual dividing lines, the consecutive head dividing lines are merged into the starting line of a table column, and the consecutive tail dividing lines are merged into the ending line of a table column; the ending lines and the starting lines of adjacent table columns are merged to generate a column line division result.
[0028] According to the present invention, a method for identifying business flow based on image recognition further includes, after generating a column and line division result:
[0029] According to the column line division result, based on the corresponding position of each column line on the X-axis, the interval between each adjacent column line is determined, and the adjacent column lines with an interval smaller than a preset threshold are divided.
[0030] According to a business flow recognition method based on image recognition provided by the present invention, based on the text in the same column and the column line division result, a numerical relationship between two columns of text in the same column is calculated to locate the balance column, including:
[0031] Constructing a basic grid unit based on the row text and the column line division result;
[0032] Convert the text in the basic grid cell that is a numeric value into a double-precision floating-point value, and calculate the sum or difference between the upper and lower rows in each two columns of values;
[0033] If the sum of the values in the jth column of the i-1th row and the kth column of the i-th row is equal to the value in the jth column of the i-th row, or if the difference between the values in the jth column of the i-1th row and the kth column of the i-th row is equal to the value in the jth column of the i-th row, the jth column is positioned as the balance column.
[0034] According to a business flow recognition method based on image recognition provided by the present invention, generating a row line division result between adjacent main rows based on the main row, including:
[0035] Determine the range of all possible separator rows between two adjacent main rows;
[0036] In each possible separator line, counting the number of times that all the texts in the same line between two adjacent main lines can be used as separators between the two main lines;
[0037] Determine the text in the same line that can be used as the separation between two main lines the greatest number of times, and use the text in the same line as the line division result between adjacent main lines.
[0038] The present invention also provides a business flow recognition device based on image recognition, comprising:
[0039] A text line recognition module is used to obtain the text line position of the original image and crop it to obtain a text line image segment, input the text line image segment into a recognition model, and obtain the recognition result, single word content and single word position of the text line output by the recognition model;
[0040] a table detection module, configured to input the original image into a table detection model and obtain a position mark of a table area output by the table detection model;
[0041] A text-in-line construction module, configured to combine and connect adjacent text lines belonging to the same line based on the position annotations to construct text-in-line;
[0042] A column division module is used to generate a column and line division result of the same text in a preset order based on the content of the single word and the position of the single word;
[0043] a row information merging module, configured to calculate, based on the text in the same row and the column line division result, a numerical relationship between the text in the same row of two columns to locate a balance column, define the text in the same row corresponding to the balance column as a main row, and generate a row line division result between adjacent main rows based on the main row;
[0044] The table information generation module is used to construct virtual row and column line coordinates based on the column line division results and the row line division results, generate virtual cells, fill the text corresponding to the main row into the corresponding virtual cells, and generate business flow results.
[0045] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the image recognition-based business flow identification method as described above is implemented.
[0046] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for business flow identification based on image recognition.
[0047] The present invention provides a method and device for business flow recognition based on image recognition, which obtains the text line position of the original image and cuts it to obtain the text line image fragment, inputs the text line image fragment into the recognition model, obtains the recognition result of the text line output by the recognition model, the single word content and the single word position; inputs the original image into the table detection model, obtains the position annotation of the table area output by the table detection model; based on the position annotation, the adjacent text lines belonging to the same row are combined and connected to construct the same text; based on the single word content and the single word position, the column line division result of the same text is generated in a preset order; based on the same text and the column line division result, the numerical relationship between the two columns of the same text is calculated to locate the balance column, the same text corresponding to the balance column is defined as the main row, and the row line division result between adjacent main rows is generated based on the main row; based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells, and the text corresponding to the main row is filled into the corresponding virtual cells to generate business flow results. Through the application of the above method and device, a business flow extraction method for wireless business tables can be constructed based on the information characteristics unique to business flow, which is fast, efficient and generalizable. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 This is a flow chart of a business flow identification method based on image recognition provided by the present invention;
[0050] Figure 2 This is one of the sample data schematic diagrams in the embodiment of the present invention;
[0051] Figure 3 This is the second schematic diagram of sample data in an embodiment of the present invention;
[0052] Figure 4 This is the third sample data schematic diagram in the embodiment of the present invention;
[0053] Figure 5 This is the fourth schematic diagram of sample data in the embodiment of the present invention;
[0054] Figure 6 This is the fifth sample data schematic diagram in the embodiment of the present invention;
[0055] Figure 7 It is a structural diagram of the business flow recognition device based on image recognition provided by the present invention;
[0056] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0057] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0058] The following combination Figures 1 to 6 The business flow recognition method based on image recognition of the present invention is described.
[0059] like Figure 1 As shown, the business flow recognition method based on image recognition provided by the present invention includes the following steps:
[0060] S1. Obtain the text line position of the original image and crop it to obtain a text line image segment, input the text line image segment into a recognition model, and obtain the recognition result, single word content and single word position of the text line output by the recognition model.
[0061] Specifically, the present invention employs a universal text detection model. An original image is input into the text detection model, and the text detection model outputs the positions of text lines in the original image. After obtaining the positions of text lines in the original image, a text line image segment is cropped from the image. The text line image segment is then input into a universal text recognition model, and the text line recognition results, individual word content, and individual word positions output by the text recognition model are obtained. The individual word positions are then mapped back to the original image.
[0062] S2. Input the original image into the table detection model to obtain the position annotation of the table area output by the table detection model.
[0063] The table detection model used in this paper is obtained through supervised training using the YOLO detection method. Specifically, various table images are collected and the images are annotated with regions. The annotations include table categories and table locations. The categories include framed tables and frameless tables. The location annotation format is the minimum bounding rectangle that can contain the table content, such as Figure 1 As shown. After the image is labeled, the labeled image is used as a sample image set, and the table detection model is trained in a supervised manner. The detection targets are divided into two types: framed tables and frameless tables. After training, the table detection model can detect the input image and obtain framed table areas and frameless table areas. Among them, the framed table has real and identifiable horizontal and vertical lines, and each cell of the table is divided by a solid line, hereinafter referred to as a wired table. The frameless table does not have complete horizontal and vertical lines, that is, its cells do not contain at least one horizontal and vertical line for division. Its cells are divided according to their content arrangement. Generally, a row of information is aligned left and right at the same height, and a column of information is aligned top and bottom, hereinafter referred to as a wireless table.
[0064] In an optional embodiment of the present invention, rotation correction is performed on the original image before inputting it into the table detection model. Because the table in the original image may appear tilted or rotated due to shooting angle, device tilt, or other reasons, rotation correction can improve the table detection model's detection performance.
[0065] S3. Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct the same line of text.
[0066] The purpose of this step is to find the text lines that belong to the same row in the table area according to the order of the text line position from left to right and from top to bottom, and connect them into the same row of text in sequence, so as to convert multiple text lines in the table into a set of the same row of text.
[0067] Specifically, step S3 includes the following sub-steps:
[0068] S301 : Sort the text lines in a preset order according to the position coordinates of the text lines.
[0069] Specifically, the text lines are sorted according to their position coordinates in order from left to right and from top to bottom, with the left to right order taking priority.
[0070] S302 , draw a foreground image (textidimg) of the text line, wherein the size of the foreground image is the same as that of the original image, the value of the text line area in the foreground image is the unique ID corresponding to each text line, and the value of other areas is 0.
[0071] For example, suppose there are two text lines in an image, with ids 1 and 2 respectively. In the foreground image corresponding to the image, the value of the area corresponding to the first text line is set to 1, the value of the area corresponding to the second text line is set to 2, and the values of the remaining areas are set to 0.
[0072] S303. According to the position mark and the sorting order of the text lines, loop to find the nearest text line belonging to the same line on the right side of each text line and connect them, and judge whether the connection meets the preset conditions based on the foreground image; concatenate the text lines that meet the preset conditions and are connected in pairs into the same line text.
[0073] Specifically, assuming that two text lines are text_i and text_j, to determine whether text lines text_i and text_j can be connected, the text lines must meet the following two conditions at the same time:
[0074] (1) Connect the center points of the left and right edges of the position annotations of the two text lines text_i, and extend the extension line to the text line text_j. When the extension line passes through the edge of the position annotation of the text line text_j that is adjacent to the text line text_i, the contact point between the extension line and the adjacent edge divides the adjacent edge into two segments, and when the ratio of the two segments is greater than a preset threshold, it is determined that the text lines text_i and text_j meet the first connection condition. Figure 2 As shown, Figure 2In the figure, four text lines, text_1, text_2, text_3, and text_4, are shown. The center points of the left and right sides of the position mark (rectangular box) of text line text_1 are connected, and the connecting line is extended to text line text_2. The extended line passes through the left side of the position mark of text line text_2 (i.e., the side adjacent to text line text_1). At this time, the left side of the position mark of text line text_2 is divided into two sections by the extended line. If the ratio of the short side length to the long side length is greater than a preset threshold (e.g., 0.6), then text lines text_1 and text_2 meet the first connection condition. Figure 2 Taking the text lines text_2 and text_3 in the example, after the left side of the position mark of the text line text_3 is divided by the extension line of the text line text2, the ratio of the short side length to the long side length is less than the preset threshold (for example, 0.6), then it is determined that the text lines text_2 and text_3 do not meet the first connection condition.
[0075] (2) Connect the center point of the position mark of text line text_i with the center point of the position mark of text line text_j. If the coordinates of the connection line always correspond to the area with a value of 0 in the foreground image, it is determined that text line text_i and text line text_j meet the second connection condition. Specifically, for two text lines that meet the first connection condition, it is also necessary to meet the requirement that the connection line of the center points of the two text lines does not pass through other text lines, that is, the coordinates of the connection line of the center points of the two text lines always correspond to the area with a value of 0 in the foreground image. Figure 2 For example, text lines text_2 and text_4 meet the first connection condition, but the line connecting their center points passes through text line text_3, and therefore does not meet the second connection condition.
[0076] If the text line text_i and the text line text_j both meet the first connection condition and the second connection condition, the text line text_i and the text line text_j are determined to meet the connection condition. If the text line text_i and the text line text_j do not meet the first connection condition and the second connection condition, the connection between the text line text_i and the text line text_j is canceled. The text lines that meet the first connection condition and the second connection condition are connected in pairs to form a text line. The connection result is as follows: Figure 3 shown.
[0077] S4. Based on the content and position of the single words, generate the column and line division results of the same text in a preset order.
[0078] In this step, for the text in the same row, segmentation candidate points are generated at the granularity of single characters in the order from left to right. In the order from top to bottom in the same row, according to the start and end attributes of the segmentation candidate points, the tangent points in the same column find the nearest segmentation candidate points of the same type downward, and the connection lines of the tangent points form virtual segmentation lines. After sorting the virtual division lines from left to right, local merging is performed to generate the final column line division result.
[0079] Specifically, step S4 includes the following sub-steps:
[0080] S401. Generate segmentation candidate points for the text line based on the single character content and single character position.
[0081] Specifically, the segmentation candidate points include two types: head tangent points and tail tangent points. The first character of the text line corresponds to the head tangent point, and the last character corresponds to the tail tangent point. In an optional embodiment of the present invention, the head tangent point and the tail tangent point are used as attributes of the single character information, and are respectively recorded by headflag and tailflag, and the initial values are both -1. When it is a head tangent point, headflag is set to 0, and when it is a tail tangent point, tailflag is set to 0. The headflag of the first character of the text line is set to 0, and the tailflag of the last character of the text line is set to 0.
[0082] When the text lines in the same row are too close to each other, it is easy to cause the result that the detected single text line is the adhesion of multiple text lines. Therefore, the segmentation candidate points are obtained by identifying the single character content and position. For example, when the distance between "AB" in the character string "ABC" exceeds 1.5 times the distance between "BC" (which can be set according to the actual situation, and this is only an example here), the tailflag of the single character "A" is set to 0, and the headflag of the single character "B" is set to 0. When the text is a transition from a numeric string to a text string, a set of head tangent points and tail tangent points will also be inserted. For example, the tailflag of the last "0" in the text string "10,000,00 Bank of China" is set to 0, and the headflag of "China" is set to 0. The schematic diagram of the tangent point candidates is as Figure 4 shown, where the rhombus is the head tangent point and the circle is the tail tangent point.
[0083] S402. According to the order from top to bottom of the text lines in the same row, find the same type of tangent points downward for the segmentation candidate points of the text line and perform virtual connection.
[0084] Specifically, according to the order from top to bottom of the text lines in the same row, find a tangent point of the same type with a horizontal distance less than a preset threshold (the threshold can be set according to the actual situation) downward for the generated head tangent point or tail tangent point for virtual connection, that is, the head tangent point looks for the head tangent point, and the tail tangent point looks for the tail tangent point. The principle of looking downward is that the vertical distance between the two tangent points to be connected is as small as possible, and the horizontal distance is as small as possible, that is, in order to find the connection of characters that are approximately left-aligned or right-aligned.
[0085] S403: Aggregate candidate segmentation points with virtual connections into a virtual segmentation line.
[0086] The virtual split lines include the headline formed by connecting the head and tail points, and the tailline formed by connecting the tail points. The headline is denoted as headline and represents the left starting line of the table column; the tailline is denoted as tailline and represents the right ending line of the table column.
[0087] S404: Calculate the credibility of the virtual dividing lines and remove the virtual dividing lines that do not meet the preset conditions.
[0088] Specifically, we first perform straight line fitting on the split points within the virtual line. We use the OpenCV line fitting function, fitline, to input a set of point coordinates and output a fitted line. We then calculate the positional deviation between each point and the fitted line. Points with a deviation of less than 6 pixels are considered small deviation points, and their number is counted. If the number of small deviation points exceeds 60%, the split line is considered reliable and retained. If the proportion of small deviation points is less than 60%, we calculate the proportion of points whose tangent points are on either side of the text. If this proportion is less than 60%, the split line is deleted.
[0089] S405: Sort the remaining virtual dividing lines in order from left to right.
[0090] Specifically, the sorting logic needs to satisfy the requirement that the virtual dividing line where the right tangent point of the same table row is located is on the right side of the virtual dividing line where the left tangent point is located.
[0091] S406: Among the sorted virtual split lines, merge the consecutive head split lines into the start line of a table column, merge the consecutive tail split lines into the end line of a table column; merge the end lines and start lines of adjacent table columns to generate a column line division result. Figure 5 shown.
[0092] In an optional embodiment of the present invention, in order to avoid the inability to generate segmentation candidate points due to dense adhesion of text lines, after step S406, it is necessary to perform secondary segmentation based on the column lines. According to the column line segmentation result, based on the corresponding position of each column line on the X-axis, the interval between each adjacent column line is determined, and the adjacent column lines with an interval less than a preset threshold are segmented. Figure 3 For example, assuming that in the row numbered "170", the contents of the "Other Party's Account Name" column and the contents of the "Summary Notes" column are stuck together, and the interval between the column lines corresponding to the "Other Party's Account Name" column and the "Summary Notes" column on the X-axis is less than the preset threshold, then the text lines of the "Other Party's Account Name" column and the text lines of the "Summary Notes" column need to be split a second time.
[0093] S5. Based on the text in the same row and the column line division results, calculate the numerical relationship between the text in the same row of two columns to locate the balance column, define the text in the same row corresponding to the balance column as the main row, and generate the row line division results between adjacent main rows based on the main row.
[0094] Among them, the method of locating the balance column is as follows:
[0095] S501: Construct basic grid cells based on the text in the same row and the column line division results.
[0096] S502: Convert the text values in the basic grid cells into double-precision floating-point values, and calculate the sum or difference of the upper and lower rows in each two columns of values. Converting to double-precision floating-point values avoids the loss of precision associated with single-precision floating-point values, ensuring the accuracy of numerical calculations.
[0097] S503. If the sum of the values in the i-1th row, jth column and the i-th row, kth column is equal to the value in the i-th row, jth column, or if the difference between the values in the i-1th row, jth column and the i-th row, kth column is equal to the value in the i-th row, jth column is positioned as the balance column. The calculation formula is as follows:
[0098] val_row(i-1)_column(j)+val_row(i)_column(k)==val_row(i)_column(j), or,
[0099] val_row(i-1)_column(j)-val_row(i)_column(k)==val_row(i)_column(j).
[0100] Among them, val_row(i)_column(j) represents the value corresponding to the i-th row and j-th column.
[0101] Generating row line division results between adjacent main rows based on the main rows includes the following steps:
[0102] S511. Determine the range of all possible separation rows between two adjacent main rows.
[0103] Specifically, if Figure 6As shown, the three lines of text surrounded by three long rectangular boxes are the main lines. In the first and second main lines, the id of the text in the same line corresponding to the first main line is recorded as i, and the id of the text in the same line corresponding to the second main line is recorded as j in order from top to bottom. The range of possible separator lines between the first and second main lines is between i+1 and j. Interval line segments A, B, C and D are possible separator candidates between the first and second main lines. Among them, the interval line segment A and the interval line segment B indicate that the range of the separator line is between i+1 and j, the interval line segment C indicates that the range of the separator line is between i+3 and j, and the interval line segment D indicates that there is only one separator line range (that is, the text line below the line segment D).
[0104] S512. In each possible separator line, count the number of times that all texts in the same line between two adjacent main lines can be used as separators between the two main lines.
[0105] S513: Determine the text in the same line that can be used as the separator between two main lines the most times, and use the text in the same line as the line division result between adjacent main lines.
[0106] Based on the interval range in step S511, the final interval indicated by multiple interval line segments can only be D interval, that is, the row "1...5" can be used as the separation between two main rows the most times.
[0107] Likewise Figure 6 For example, the area between the second and third main lines illustrates a difficult scenario. The spacing segment H is a long interval, but the upper line of the separation range indicated by the spacing segment G is the text line above the spacing segment G ("3004, AU"). The number of lines at the lower end of the spacing segment H is less than the text line above the spacing segment G. Therefore, it conflicts with the candidate range generated by the spacing segment G, and the spacing segment H candidate cannot be used due to the conflict. At this time, the available candidates are accumulated, and the text line below the spacing segment K has the largest number of cumulative values that can be used as the separation between the two main lines. The final voting result is the text below the spacing segment K. Comprehensive judgment based on the separation of each column can reduce the errors caused by simply using the maximum spacing division.
[0108] S6. Construct virtual row and column line coordinates based on the column line division results and the row line division results, generate virtual cells, fill the text corresponding to the main row into the corresponding virtual cells, and generate business flow results.
[0109] Specifically, after column division and row merging are implemented, virtual row and column coordinates can be constructed based on the virtual column and row division results at a single-word granularity to generate virtual cell information. The text corresponding to the primary row is then filled in and assembled to generate the final wireless service flow results. In an optional embodiment of the present invention, the wireless service flow is used as part of flow recognition, combined with the results of the wired service flow and text outside the table, ultimately achieving text information and structure extraction for the entire image page.
[0110] In summary, the business flow recognition method based on image recognition provided by the present invention constructs a text table structure through text recognition information. The introduction of single-word content and single-word position information improves the ability to divide table columns with close distances or nearly glued single words, so that the method has good generalization. After the same-line combination and column division are performed, the position of the amount column is inferred through the flow amount calculation unique to the business flow. The stable amount column position can be used as an important reference for merging multiple lines of text into a single information. After constructing the basic horizontal and vertical grid structure, the division candidates are generated according to the upper and lower intervals of the multi-column cells of the same line and combined voting is performed. The combination method of multiple information rows generated by voting is stable, which can reduce the erroneous division caused by the single longest interval and stabilize the extraction results of complex wireless bank flow. The business flow recognition method provided by the present invention is based on the information characteristics unique to the business flow, and is a method for extracting wireless business flow. It is fast, efficient, generalizable, and has good application prospects.
[0111] Based on the same inventive concept, the present invention also provides a business flow identification device based on image recognition. The business flow identification device based on image recognition provided by the present invention is described below. The business flow identification device based on image recognition described below and the business flow identification method based on image recognition described above can be referenced to each other.
[0112] like Figure 7 As shown, the business flow recognition device based on image recognition provided by the present invention includes a text row recognition module 701, a table detection module 702, a peer text construction module 703, a column division module 704, a row information merging module 705 and a table information generation module 706.
[0113] The text line recognition module 701 is used to obtain the text line position of the original image and crop it to obtain the text line image segment, input the text line image segment into the recognition model, and obtain the text line recognition result, single word content and single word position output by the recognition model.
[0114] The table detection module 702 is used to input the original image into the table detection model and obtain the position annotation of the table area output by the table detection model.
[0115] The same-line text construction module 703 is used to combine and connect adjacent text lines belonging to the same line to construct the same-line text based on position annotations.
[0116] The column division module 704 is used to generate column and line division results of the same line of text based on the content and position of the single words in a preset order.
[0117] The row information merging module 705 is used to calculate the numerical relationship between the text in the same row of two columns based on the text in the same row and the column line division results to locate the balance column, define the text in the same row corresponding to the balance column as the main row, and generate the row line division results between adjacent main rows based on the main row.
[0118] The table information generation module 706 is used to construct virtual row and column line coordinates based on the column line division results and the row line division results, generate virtual cells, fill the text corresponding to the main row into the corresponding virtual cells, and generate business flow results.
[0119] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the business flow recognition method based on image recognition provided by the above methods, which includes:
[0120] Obtaining the text line position of the original image and cropping it to obtain a text line image segment, inputting the text line image segment into a recognition model, and obtaining the recognition result, single word content, and single word position of the text line output by the recognition model;
[0121] Inputting the original image into a table detection model to obtain a position mark of the table area output by the table detection model;
[0122] Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line;
[0123] Based on the individual word content and the individual word position, generating the column line division result of the text in the same row according to a preset order;
[0124] Based on the text in the same row and the column line division result, a numerical relationship between two columns of text in the same row is calculated to locate a balance column, the text in the same row corresponding to the balance column is defined as a main row, and a row line division result between adjacent main rows is generated based on the main row;
[0125] Based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells. The text corresponding to the main row is filled into the corresponding virtual cells to generate a business flow result.
[0126] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0127] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the business flow identification method based on image recognition provided by the above methods, which method comprises:
[0128] Obtaining the text line position of the original image and cropping it to obtain a text line image segment, inputting the text line image segment into a recognition model, and obtaining the recognition result, single word content, and single word position of the text line output by the recognition model;
[0129] Inputting the original image into a table detection model to obtain a position mark of the table area output by the table detection model;
[0130] Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line;
[0131] Based on the individual word content and the individual word position, generating the column line division result of the text in the same row according to a preset order;
[0132] Based on the text in the same row and the column line division result, a numerical relationship between two columns of text in the same row is calculated to locate a balance column, the text in the same row corresponding to the balance column is defined as a main row, and a row line division result between adjacent main rows is generated based on the main row;
[0133] Based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells. The text corresponding to the main row is filled into the corresponding virtual cells to generate a business flow result.
[0134] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the business flow identification method based on image recognition provided by the above methods, the method comprising:
[0135] Obtaining the text line position of the original image and cropping it to obtain a text line image segment, inputting the text line image segment into a recognition model, and obtaining the recognition result, single word content, and single word position of the text line output by the recognition model;
[0136] Inputting the original image into a table detection model to obtain a position mark of the table area output by the table detection model;
[0137] Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line;
[0138] Based on the individual word content and the individual word position, generating the column line division result of the text in the same row according to a preset order;
[0139] Based on the text in the same row and the column line division result, a numerical relationship between two columns of text in the same row is calculated to locate a balance column, the text in the same row corresponding to the balance column is defined as a main row, and a row line division result between adjacent main rows is generated based on the main row;
[0140] Based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells. The text corresponding to the main row is filled into the corresponding virtual cells to generate a business flow result.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A business flow recognition method based on image recognition, characterized in that: include: Obtaining the text line position of the original image and cropping it to obtain a text line image segment, inputting the text line image segment into a recognition model, and obtaining the recognition result, single word content, and single word position of the text line output by the recognition model; Inputting the original image into a table detection model to obtain a position mark of the table area output by the table detection model; Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line; Based on the individual word content and the individual word position, generating the column line division result of the text in the same row according to a preset order; Based on the text in the same row and the column line division result, a numerical relationship between two columns of text in the same row is calculated to locate a balance column, the text in the same row corresponding to the balance column is defined as a main row, and a row line division result between adjacent main rows is generated based on the main row; Based on the column line division result and the row line division result, virtual row and column line coordinates are constructed to generate virtual cells. The text corresponding to the main row is filled into the corresponding virtual cells to generate a business flow result.
2. The business flow recognition method based on image recognition according to claim 1, characterized in that: Based on the position annotation, adjacent text lines belonging to the same line are combined and connected to construct text in the same line, including: sorting the text lines in a preset order according to the position coordinates of the text lines; Draw a foreground image of the text line, wherein the size of the foreground image is the same as that of the original image, the value of the text line area in the foreground image is a unique ID corresponding to each text line, and the value of other areas is 0; According to the position mark and the sorting order of the text lines, loopingly find the nearest text line belonging to the same line on the right side of each text line and connecting them, and judging whether the connection meets the preset conditions based on the foreground image; The text lines connected in pairs that meet the preset conditions are concatenated into the same line of text.
3. The business flow recognition method based on image recognition according to claim 2, characterized in that: According to the position mark and the sorting order of the text lines, loopingly finding the nearest text line belonging to the same line on the right side of each text line and connecting them, and judging whether the connection meets a preset condition based on the foreground image, including: Connect the center points of the left side and the right side of the position mark of text line text_i, and extend the extension line to text line text_j. When the extension line passes through the edge of the position mark of text line text_j that is adjacent to text line text_i, and the contact point between the extension line and the adjacent edge divides the adjacent edge into two segments, and the ratio of the two segments is greater than a preset threshold, it is determined that text line text_i and text line text_j meet the first connection condition. Connecting a center point of the position mark of text line text_i with a center point of the position mark of text line text_j, and determining that text line text_i and text line text_j meet a second connection condition when the coordinates of the connecting line always correspond to an area with a value of 0 in the foreground image; In the case that the text line text_i and the text line text_j satisfy both the first connection condition and the second connection condition, it is determined that the text line text_i and the text line text_j satisfy the connection condition.
4. The business flow recognition method based on image recognition according to claim 1, characterized in that: Based on the single word content and the single word position, generating the column and line division result of the text in the same row in a preset order, including: Based on the word content and the word position, generating candidate segmentation points for the text line, the candidate segmentation points including a head segmentation point and a tail segmentation point, the first word of the text line corresponding to the head segmentation point, and the last word corresponding to the tail segmentation point; According to the order of the text lines in the same row from top to bottom, searching downward for similar cut points for the candidate cut points of the text lines and performing virtual connections; Aggregating the candidate segmentation points with virtual connections into a virtual segmentation line; the virtual segmentation line includes a head segmentation line formed by a line connecting the head segmentation points and a tail segmentation line formed by a line connecting the tail segmentation points; Performing credibility calculation on the virtual dividing lines, and removing the virtual dividing lines that do not meet preset conditions; Sort the remaining virtual dividing lines in order from left to right; Among the sorted virtual dividing lines, the consecutive head dividing lines are merged into the starting line of a table column, and the consecutive tail dividing lines are merged into the ending line of a table column; the ending lines and the starting lines of adjacent table columns are merged to generate a column line division result.
5. The business flow recognition method based on image recognition according to claim 4 is characterized in that: After generating the column line division results, it also includes: According to the column line division result, based on the corresponding position of each column line on the X-axis, the interval between each adjacent column line is determined, and the adjacent column lines with an interval smaller than a preset threshold are divided.
6. The business flow recognition method based on image recognition according to claim 1, characterized in that: Calculating a numerical relationship between two columns of text in the same row based on the text in the same row and the column line division result to locate the balance column includes: Constructing a basic grid unit based on the row text and the column line division result; Convert the text in the basic grid cell that is a numeric value into a double-precision floating-point value, and calculate the sum or difference between the upper and lower rows in each two columns of values; If the sum of the values in the jth column of the i-1th row and the kth column of the i-th row is equal to the value in the jth column of the i-th row, or if the difference between the values in the jth column of the i-1th row and the kth column of the i-th row is equal to the value in the jth column of the i-th row, the jth column is positioned as the balance column.
7. The method for identifying business flow based on image recognition according to any one of claims 1 to 6, characterized in that: Generating row line division results between adjacent main rows based on the main rows includes: Determine the range of all possible separator rows between two adjacent main rows; In each possible separator line, counting the number of times that all the texts in the same line between two adjacent main lines can be used as separators between the two main lines; Determine the text in the same line that can be used as the separation between two main lines the greatest number of times, and use the text in the same line as the line division result between adjacent main lines.
8. A business flow recognition device based on image recognition, characterized in that: include: A text line recognition module is used to obtain the text line position of the original image and crop it to obtain a text line image segment, input the text line image segment into a recognition model, and obtain the recognition result, single word content and single word position of the text line output by the recognition model; a table detection module, configured to input the original image into a table detection model and obtain a position mark of a table area output by the table detection model; A text-in-line construction module, configured to combine and connect adjacent text lines belonging to the same line based on the position annotations to construct text-in-line; A column division module is used to generate a column and line division result of the same text in a preset order based on the content of the single word and the position of the single word; a row information merging module, configured to calculate, based on the text in the same row and the column line division result, a numerical relationship between the text in the same row of two columns to locate a balance column, define the text in the same row corresponding to the balance column as a main row, and generate a row line division result between adjacent main rows based on the main row; The table information generation module is used to construct virtual row and column line coordinates based on the column line division results and the row line division results, generate virtual cells, fill the text corresponding to the main row into the corresponding virtual cells, and generate business flow results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the business flow identification method based on image recognition is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the business flow identification method based on image recognition as described in any one of claims 1 to 7 is implemented.