Finance and tax bill classification detection system based on image recognition
By designing a fiscal and tax bill classification detection system based on image recognition, differentiated optimization is carried out for different areas in the bill image, the problem of poor image enhancement effect in the prior art is solved, and the precise classification and recognition of bills is realized.
Patent Information
- Application Number
- CN202510621677.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art is difficult to differentiately optimize differentially for different areas in the bill image, resulting in poor image enhancement effect and affecting the accuracy of bill classification results.
A fiscal and tax bill classification and detection system based on image recognition is designed, and differentiated optimization of bill images is achieved through the image acquisition module, the expansion area determination module, the expansion area classification module and the detection module. The specific steps include grayscale processing, edge detection and expansion operations on the initial image, calculating the effective probability of the expansion area, and dividing the effective area and the invalid area according to the probability, performing gamma transformation and adaptive filtering, and finally inputting the pre-trained classification detection model for ticket type identification.
By accurately distinguishing fields and defective areas in the bill image, the contrast of the effective area is enhanced, and the invalid information is effectively filtered, the image enhancement effect is significantly improved, thereby achieving accurate classification and recognition of bills.
Smart Images

Figure CN120148058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology. More specifically, the present invention relates to a financial and tax bill classification and detection system based on image recognition. Background Art
[0002] In the digital economy era, the digital management of financial and tax bills has become a core link for enterprises and government departments to improve operational efficiency and optimize resource allocation. With the in-depth promotion of digital transformation, enterprises need to efficiently classify, detect, and recognize a large number of financial and tax bills to support financial accounting, tax compliance, and intelligent decision-making. However, the key prerequisite for digital bill processing lies in obtaining high-quality image data, and the original bills often suffer from image quality degradation due to the acquisition environment (such as uneven illumination, equipment noise) or physical state (such as stains, wrinkles), specifically manifested as defects such as blurred text, background noise interference, and insufficient local contrast.
[0003] Currently, the preprocessing technology for bill images mainly relies on traditional filtering algorithms (such as mean filtering, Gaussian filtering) and contrast enhancement methods (such as gamma transformation, histogram equalization), etc. However, when traditional filtering algorithms suppress noise, they are prone to blurring the edges of key bill information (such as small-sized text, numerical symbols), reducing the accuracy of subsequent optical character recognition (OCR); although enhancement techniques such as gamma transformation can improve the distinguishability between text and background, they will simultaneously amplify the invalid information (such as stains, scratches) on the bill surface, and even introduce artifacts, exacerbating the semantic noise of the image.
[0004] Existing methods mostly adopt fixed parameters or global operations, and it is difficult to perform differential optimization for the characteristics of different regions in the bill image (such as text-dense regions and blank background regions), resulting in local imbalance in the enhancement effect. Summary of the Invention
[0005] In order to solve the problem that it is impossible to perform differential optimization for different regions in the bill image, resulting in poor enhancement effect of the bill image and thus affecting the accuracy of the bill classification result, the present invention provides a financial and tax bill classification and detection system based on image recognition. The system includes: An image acquisition module, configured to acquire the front image of the bill to be classified and perform grayscale processing to obtain an initial image; An expansion region determination module, configured to perform edge detection on the initial image and perform a horizontal expansion operation on the non-linear edges in the detection result to obtain the expansion regions of each non-linear edge; An expansion region classification module is used to select a target expansion region, calculate the effective probability of the target expansion region based on various difference features between the field edge and the defect edge, and divide all expansion regions into effective regions and invalid regions by comparing the effective probability with a preset value; wherein, the effective probability represents the probability that the target expansion region is the region where the field edge is located. A detection module is used to increase the contrast of the gray level inside each effective region through gamma transformation based on the effective probability of each effective region, and adaptively adjust the filtering window to perform mean filtering on each invalid region. The size of the filtering window is negatively correlated with the effective probability of the corresponding invalid region, so as to input the obtained target image into a pre-trained classification and detection model, and identify the type of bill according to the output text region and content.
[0006] The present invention can accurately distinguish the region where the field is located and the region where the defect is located in the bill image, thereby providing accurate data support for the differential optimization of different regions in the bill image. Moreover, by increasing the contrast of the gray level inside each effective region, the contrast between the effective information and other information can be increased. By adaptively adjusting the filtering window to perform mean filtering on each invalid region, the invalid information can be effectively filtered, improving the image enhancement effect. Thus, based on the enhanced target image, the accurate classification of the bill image can be realized.
[0007] Preferably, the method for obtaining the effective probability of the target expansion region includes: Calculating a first index of the target expansion region : ; In the formula, is the area of the inner edge of the target expansion region; is the area of the target expansion region; is the median of the lengths of all edges in the target expansion region; is the width of the circumscribed rectangle of the target expansion region; is the information entropy of the target expansion region; Extend the width of the circumscribed rectangle of the target expansion region from both ends, use the newly added region as the expansion region, calculate a second index of the target expansion region. The second index is positively correlated with the difference in information entropy between the target expansion region and the expansion region, and negatively correlated with the deviation degree of the width of the circumscribed rectangle from the overall level. Take the mean of the first index and the second index as the effective probability of the target expansion region.
[0008] The present invention synthesizes various difference features of the field edge and the defect edge, calculates the effective probability of each expansion region, and ensures the accuracy of the calculation result.
[0009] Preferably, the method for obtaining the second index of the target expansion region includes: Calculate the difference between the width of the circumscribed rectangle of the target inflated region and the median of the widths of the circumscribed rectangles of all inflated regions, and normalize it to obtain the degree of deviation of the width of the circumscribed rectangle of the target inflated region from the overall level; Calculate the difference between 1 and the degree of deviation, and perform a multiplication operation on the difference and the difference in information entropy between the target inflated region and the extended region to obtain the second index of the target inflated region.
[0010] Preferably, the method for obtaining the information entropy of the target inflated region or the extended region includes: Divide the target inflated region or the extended region into a number of square sub-regions with a set width, and calculate the information entropy of the target inflated region or the extended region based on the probabilities of different values of the number of inner edge pixel points in all sub-regions using the information entropy calculation formula.
[0011] Preferably, based on the effective probabilities of each effective region, increasing the contrast of the gray levels inside each effective region through gamma transformation includes: Use the sum of the effective probability of each effective region and 1 as the weight to weight the preset gamma coefficient to obtain the updated gamma coefficient for each effective region; Use the updated gamma coefficient to adjust the gray values of the pixel points inside each effective region through gamma transformation.
[0012] The present invention can greatly enhance the contrast of the gray levels inside the effective regions with relatively large effective probabilities, thereby ensuring the highest contrast between the fields and other information.
[0013] Preferably, the adaptive adjustment of the filtering window satisfies the following relational expression: ; In the formula, is the width of the adaptive filtering window of the th invalid region; is the effective probability of the th invalid region; is the width of the preset initial filtering window; is the preset magnification factor; is the ceiling function.
[0014] The present invention sets a relatively large filtering window for the invalid regions with relatively low effective probabilities, which can greatly smooth the invalid information, thereby reducing the interference of the invalid information.
[0015] Preferably, the method for obtaining the horizontal direction includes: Perform Hough line detection on the initial image, and cluster all the detected line edges with lengths greater than the preset length based on the slope, so as to take the line direction corresponding to the cluster center of the category with a smaller average slope as the horizontal direction, and take the line direction corresponding to the cluster center of the other category as the vertical direction.
[0016] Preferably, the method for obtaining non-line edges includes: Obtain all the edges obtained by performing edge detection on the initial image and all the line edges obtained by performing Hough line detection on the initial image; Define all the edges other than the line edges among all the edges as non-line edges.
[0017] The purpose of screening non-line edges in the present invention is to eliminate the invalid information of the boundary lines in the bill image, thereby improving the efficiency of edge analysis.
[0018] Preferably, the length direction of the circumscribed rectangle of the target dilation region is parallel to the horizontal direction, and the width direction is parallel to the vertical direction.
[0019] Preferably, the classification detection model is an algorithm model combining a target detection model of YOLOv10 and an OCR recognition model.
[0020] The present invention has the following effects: The present invention uses the distribution characteristics of fields (that is, fields are usually distributed row by row in the bill image) to perform horizontal dilation operations on each non-line edge, which can ensure finding all fields. Then, using various difference characteristics between field edges and defect edges, calculate the probability that each dilation region is a field edge, and distinguish the valid region and the invalid region, which can realize the distinction between the field region and the defect region. Then, increase the contrast of the internal gray level of the valid region, which can increase the contrast between the fields and other information, and perform mean filtering on each invalid region, which can effectively smooth the invalid information, thereby improving the enhancement effect of the initial image, providing an accurate image basis for the subsequent recognition of bill types, and further realizing the accurate classification of bills. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become easily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein: Figure 1 It is a schematic structural diagram of a fiscal and tax bill classification detection system based on image recognition according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] The following will describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings.
[0024] Refer to Figure 1 , a fiscal and tax bill classification and detection system 100 based on image recognition includes an acquisition module 110, a dilation region determination module 120, a dilation region classification module 130, and a detection module 140, which are specifically as follows: The acquisition module 110 is configured to acquire the front image of the bill to be classified and perform grayscale processing to obtain an initial image.
[0025] Specifically, a high-resolution scanner (recommended ≥300 dpi), a digital camera, or a mobile phone shooting device can be used to shoot the front image of the bill to be classified tiled on a solid color background (such as white), requiring the four corners of the bill to be aligned with the edges of the picture, and by the weighted average method or the average value method, the RGB three channels of the collected color image are merged into a single-channel grayscale image to obtain an initial image.
[0026] It should be noted that the process of grayscale processing by the weighted average method and the average value method is a prior art, and this embodiment will not elaborate on it here.
[0027] The dilation region determination module 120 is configured to perform edge detection on the initial image and perform a horizontal dilation operation on the non-linear edges in the detection result to obtain the dilation regions of each non-linear edge.
[0028] It should be noted that the information in the bill often appears as edges in the bill image. By performing edge detection on the initial image, all the information in the bill can be obtained. And the straight edges are usually the boundary lines in the bill. Therefore, the present invention only analyzes the non-linear edges, which can improve the edge analysis efficiency while extracting the key text information in the bill image.
[0029] It should be further noted that since the text information in the bill is usually horizontally distributed line by line, and the distance between words in the same field is small, and each field can be enclosed by a rectangular box, while the invalid information (such as stains, scratches, etc.) in the bill does not have the above characteristics. Therefore, in the present invention, by performing a horizontal dilation operation on each non-linear edge, the same field can be enclosed by a rectangular box. Then, based on the difference characteristics between the field edge and the defect edge, each rectangular box can be classified to accurately distinguish the area where the field edge is located and the area where the defect edge is located, providing a data basis for subsequent operations.
[0030] In an exemplary embodiment of the present invention, the determination of the horizontal direction can be achieved through the following steps: Perform Hough line detection on the initial image, and cluster all the detected line edges with a length greater than a preset length based on the slope, so as to take the line direction corresponding to the cluster center of the category with a smaller average slope as the horizontal direction, and take the line direction corresponding to the cluster center of the other category as the vertical direction.
[0031] It should be noted that since the slope of the horizontal direction is zero and the slope of the vertical direction is close to 90 degrees. Therefore, taking the line direction corresponding to the cluster center of the category with a smaller average slope as the horizontal direction of the initial image and taking the line direction corresponding to the cluster center of the other category as the vertical direction of the initial image can achieve accurate discrimination of the plane direction. Among them, the process of Hough line detection is a prior art, and this embodiment will not elaborate here. Exemplarily, the preset length can be set to ; where is the width of the initial image (i.e., the length value in the vertical direction).
[0032] In an exemplary embodiment of the present invention, the determination of the non-linear edges in the initial image can be achieved through the following steps: Obtain all the edges obtained by performing edge detection on the initial image and all the line edges obtained by performing Hough line detection on the initial image; define all the edges except each line edge among all the edges as non-linear edges.
[0033] Exemplarily, the Canny edge detection algorithm, Sobel edge detection, etc. can be used to perform edge detection on the initial image to obtain all the edges in the initial image, and then take the edges except all the line edges detected by Hough line detection among the obtained all the edges as non-linear edges, so as to eliminate the edges of the boundary lines in the bill image, reduce the number of edge analyses, and improve the analysis efficiency of subsequent operations.
[0034] Next, the process of determining the dilation regions of each non-linear edge will be described in detail: First, select a template of appropriate size (also known as a structuring element, a professional term in image dilation), such as a rectangle with a length of 3 and a width of 1, and set the number of iterations of the dilation operation according to an empirical value (such as 5 times); then, use the template to scan each pixel point in any non-linear edge. If at least one pixel value in the area covered by the template is 1, set the pixel value of the corresponding pixel point to 1, so as to expand the highlighted area, achieve the dilation effect, and so on until the set number of iterations is reached, obtaining the dilation regions of each non-linear edge.
[0035] It should be noted that since the present invention only needs to perform dilation operations on each non-linear edge in the horizontal direction, the width of the template is set to 1 in the present invention.
[0036] The dilation region classification module 130 is used to select a target dilation region, calculate the effective probability of the target dilation region based on various difference features between the field edge and the defect edge, and divide all dilation regions into effective regions and invalid regions by comparing the effective probability with a preset value; where the effective probability represents the probability that the target dilation region is the region where the field edge is located.
[0037] Among them, the target dilation region refers to a dilation region randomly selected from all the dilation regions determined by module 120.
[0038] In an exemplary embodiment of the present invention, the determination of the effective probability of the target dilation region can be achieved through the following steps: Step 1: Calculate the first index of the target dilation region; Among them, the first index is a parameter used to measure the possibility that the target dilation region is the region where the field edge is located.
[0039] Specifically, the first index of the target dilation region satisfies the following relational expression: ; In the formula, is the first index of the target dilation region; is the area of the inner edge of the target dilation region; is the area of the target dilation region; is the median of the lengths of all edges in the target dilation region; is the width of the circumscribed rectangle of the target dilation region; is the information entropy of the target dilation region, and this value reflects the degree of chaos of the distribution of inner edge pixel points in the target dilation region.
[0040] Among them, the area of the inner edge of the target dilation region can be represented by the number of all edge pixel points in the target dilation region.
[0041] It reflects the density of edges within the target dilated region. The larger this value is, the denser the edge distribution within the target dilated region, and thus the greater the likelihood that the target dilated region is the region where the field edges are located, and the corresponding first index is relatively large.
[0042] It should be noted that since the longest edge length in the text is usually 4 times the width of the text, while the length of the defect edges is relatively random. Therefore, the present invention utilizes this feature to calculate the value of. The larger this value is, the more the edges within the target dilated region conform to the characteristics of the text edges, and thus the greater the likelihood that the target dilated region is the region where the field edges are located, and the corresponding first index is relatively large. Among them, the 4 times here is a general rule found when extracting the text edges in the intact bill image. Of course, according to specific situations, the ratio of the longest edge of the text in the corresponding image to the width of the corresponding text can also be calculated to determine the number in the relational expression composed of and in the first index.
[0043] It should be further noted that the edges of the text or numbers in the field are usually relatively complex and rich, which makes the distribution of the field edges relatively chaotic. And for the defects in the bill image, such as stains, scratches, patches, etc., compared with the text or numbers, the number of their edges is relatively small and the degree of chaos is relatively small. And information entropy is a data that can evaluate the degree of chaos of information distribution. Therefore, when the information entropy of the target dilated region is relatively large, it can be explained that the likelihood that the target dilated region is the region where the field edges are located is relatively large, and the corresponding first index is relatively large.
[0044] In an exemplary embodiment of the present invention, the determination of the information entropy of the target dilated region can be achieved through the following steps: Divide the target dilated region into a number of square sub-regions with a set width, and based on the probabilities of different values of the number of edge pixel points within each sub-region, use the calculation formula of information entropy to calculate the information entropy of the target dilated region.
[0045] Exemplarily, the target dilated region can be divided into a number of square sub-regions with a width of 4, and then the information entropy of the target dilated region is calculated. Specifically, the information entropy of the target dilated region satisfies the following relational expression: ; In the formula, is the information entropy of the target dilated region; is the probability that the number of edge pixel points within all sub-regions of the target dilated region is ; is the logarithmic function; is the summation symbol.
[0046] It should be noted that since the target dilation region in the present invention is divided into a number of square sub-regions with a width of 4, that is, each sub-region contains 16 pixel points. Therefore, when calculating the information entropy of the target dilation region, the 16 in the calculation formula is the number of pixel points in the sub-region.
[0047] Step 2: Extend the width of the circumscribed rectangle of the target dilation region from both ends, and use the newly added region as the extended region. Calculate the second index of the target dilation region. The second index is positively correlated with the difference in information entropy between the target dilation region and the extended region, and negatively correlated with the degree of deviation of the width of the circumscribed rectangle from the overall horizontal level.
[0048] Among them, the second index refers to a parameter that can measure the possibility that the target dilation region is the region where the field edge is located by combining the surrounding regions.
[0049] In an exemplary embodiment of the present invention, the length direction of the circumscribed rectangle of the target dilation region is parallel to the horizontal direction, and the width direction is parallel to the vertical direction.
[0050] Next, the determination process of the extended region will be described: Keep the centroid and length of the circumscribed rectangle of the target dilation region unchanged, and increase the width of the circumscribed rectangle by 0.2 times of the width on both sides respectively, so as to obtain a rectangular region with an unchanged length and a width 1.4 times the original width. And the region in the new rectangular region except the original circumscribed rectangle is used as the extended region.
[0051] It should be noted that since the text information in the bill is usually distributed horizontally row by row, that is, there are large blank areas in the vertical direction between fields, making the difference in information entropy between the region where the field edge is located and the extended region relatively large, while the position where the defect edge exists is random, resulting in a relatively small difference in information entropy between the region where the defect edge is located and the extended region. Therefore, the present invention uses this difference feature to measure the possibility that the target dilation region is the region where the field edge is located by calculating the difference in information entropy between the original circumscribed rectangle and the extended region.
[0052] In an exemplary embodiment of the present invention, the determination of the second index of the target dilation region can be achieved through the following steps: (1) Calculate the difference between the width of the circumscribed rectangle of the target dilation region and the median of the widths of the circumscribed rectangles of all dilation regions, and normalize it to obtain the degree of deviation of the width of the circumscribed rectangle of the target dilation region from the overall horizontal level; Optionally, the average width of the circumscribed rectangles of all dilation regions can also be defined as the overall level of the widths of the circumscribed rectangles of all dilation regions, so as to obtain this degree of deviation.
[0053] Optionally, normalization can be performed using the maximum and minimum values, or a function can be used for normalization. This embodiment does not make a special limitation on the selected normalization method.
[0054] (2) Calculate the difference between 1 and the degree of deviation, and perform a multiplication operation on the difference and the difference in information entropy between the target expansion region and the expansion region to obtain the second index of the target expansion region.
[0055] Specifically, the second index of the target expansion region satisfies the following relational expression: ; In the formula, is the second index of the target expansion region; is the width of the circumscribed rectangle of the target expansion region; is the median of the widths of the circumscribed rectangles of all expansion regions; 、 are the maximum and minimum values of the widths of the circumscribed rectangles of all expansion regions respectively; is the absolute value symbol; 、 are the information entropies of the target expansion region and the expansion region respectively.
[0056] Among them, reflects the degree of deviation of the width of the circumscribed rectangle of the target expansion region from the overall level. The larger this value is, the more inconsistent the width of the circumscribed rectangle of the target expansion region is with the overall level. Since the widths of all fields in the bill image are highly consistent, the larger this value is, the lower the possibility that the target expansion region is the region where the field edge is located, and the corresponding second index is smaller.
[0057] In another embodiment, the formula can also be used to calculate the second index of the target expansion region; in the formula, is the exponential function with the natural constant as the base.
[0058] In an exemplary embodiment of the present invention, the method for determining the information entropy of the expansion region is the same as that of the target expansion region, and this embodiment will not be elaborated here.
[0059] Step three: Use the mean of the first index and the second index as the effective probability of the target expansion region.
[0060] Specifically, the effective probability of the target expansion region satisfies the relational expression ; in the formula, 、 are the first index and the second index of the target expansion region respectively. When When it is, the target dilation region is divided into valid regions, that is, the regions where the field edges are located. Otherwise, the target dilation region is divided into invalid regions, so that the division results of all dilation regions can be obtained, providing accurate data support for subsequent operations. In this embodiment .
[0061] The detection module 140 is configured to increase the contrast of the gray levels inside each valid region through gamma transformation based on the valid probabilities of the respective valid regions, and adaptively adjust the filtering window to perform mean filtering on each invalid region. The size of the filtering window is negatively correlated with the valid probability of the corresponding invalid region, so as to input the obtained target image into a pre-trained classification and detection model, and identify the type of the bill according to the output text region and content.
[0062] In an exemplary embodiment of the present invention, the enhancement of the contrast of the gray levels inside each valid region can be achieved through the following steps: Use the sum of the valid probability of each valid region and 1 as the weight to weight the preset gamma coefficient to obtain the updated gamma coefficient for each valid region; use the updated gamma coefficient to adjust the gray level values of the pixel points inside each valid region through gamma transformation.
[0063] Specifically, the updated gray level value of any pixel point inside any valid region satisfies the following relational expression: ; In the formula, , are respectively the gray level values of the th pixel point before and after update inside the th valid region; is the valid probability of the th valid region; is the preset initial gamma coefficient, and in this embodiment .
[0064] Among them, is to normalize the gray level value of the th pixel point inside the th valid region to ensure that this value is less than 1, so that when the gamma coefficient is larger, the gray level value of this pixel point can be reduced to a greater extent.
[0065] It should be noted that since the grayscale value of the field in the bill image is relatively low, the gamma coefficient needs to be greater than 1. Moreover, the greater the gamma coefficient, the smaller the grayscale value of the updated field, that is, the better the enhancement effect on the field. Therefore, in the present invention, the initial gamma coefficient is updated by the sum of the effective probability of each effective region and 1, so that the effective regions with a greater possibility of being the region where the field is located have a greater gamma coefficient, thereby enabling a greater increase in the contrast between the field part and other parts within the corresponding effective region, and realizing the enhancement of the contrast of the grayscale within each effective region.
[0066] It should be further noted that for each invalid region, when the effective probability of any invalid region is relatively low, it indicates that the possibility of this invalid region being the region where the field edge is located is extremely small, that is, this invalid region is very likely to be the region where the defect edge is located. By setting a larger filtering window, the edge within this invalid region can be better smoothed, thereby improving the filtering effect of the defect edge.
[0067] Specifically, the size of the adaptive filtering window satisfies the following relational expression: ; In the formula, is the width of the adaptive filtering window of the th invalid region; is the effective probability of the th invalid region; is the width of the preset initial filtering window, in this embodiment ; is the preset magnification factor, in this embodiment ; is the ceiling function.
[0068] In another embodiment, the size of the adaptive filtering window for each invalid region can also be calculated by the formula: :
[0069] Among them, when the contrast of the grayscale within all effective regions is increased and the filtering of all invalid regions is completed, the obtained image is the target image.
[0070] Furthermore, the target image can be input into a pre-trained classification and detection model, and the position of the text region and the corresponding text content are output, thereby realizing the classification of the bill image and the recognition of the corresponding text content.
[0071] In an exemplary embodiment of the present invention, the classification and detection model is an algorithm model that combines the object detection model of YOLOv10 and the OCR recognition model.
[0072] It should be noted that the object detection model of YOLOv10 can find the bounding boxes of each text region in the target image, determine which category in the labels set during model training each bounding box belongs to, and then be able to crop out the classified regions; the OCR recognition model can recognize the text content in each cropped region, so that the trained classification and detection model can recognize the positions of the text regions in the target image and the corresponding text content, and based on the recognized text content, classify each bill image.
[0073] Next, the construction process of the training set of the classification and detection model will be described in detail: First, use the processed images corresponding to the front images of multiple bills (the same processing operations as for determining the target image) to construct a sample image set; then, use annotation tools (such as LabelImg, CVAT) to select the target text regions (such as "amount", "invoice number") in each image in the sample image set, and assign 6 preset labels to them: taxpayer identification number, project name, purchaser name, quantity, amount, invoice number; then, divide the annotated sample image set into a training set, a validation set, and a test set according to a ratio (such as 7:2:1) to train the classification and detection model. The labels set in this embodiment are related to the bill types, and relevant personnel can also set appropriate labels according to specific situations.
[0074] It should be noted that the process of training the model based on the training set, validation set, and test set is prior art, and this embodiment will not elaborate on it here.
[0075] Optionally, the object detection model in the classification and detection model can also select other algorithms in the YOLO series, such as YOLOv8, YOLOv9, etc.; the classification and detection model can also adopt an end-to-end text detection and recognition model, such as MaskTextSpotter, etc.; of course, an appropriate model can also be selected according to specific situations, and this embodiment does not make special limitations on this.
[0076] In the description of this specification, the meanings of "a plurality of" and "several" are at least two, for example, two, three or more, etc., unless otherwise clearly and specifically defined.
[0077] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and concept of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted during the practice of the present invention.
Claims
1. The tax bill classification detection system based on image recognition is characterized by: include: An image acquisition module is used to acquire the front image of the bill to be classified and perform grayscale processing to obtain an initial image; An expansion region determination module, used for performing edge detection on the initial image, and performing a horizontal expansion operation on the non-straight edges in the detection result to obtain the expansion region of each non-straight edge; The expansion region classification module is used to select the target expansion region, calculate the effective probability of the target expansion region based on the multiple difference characteristics between the field edge and the defect edge, and divide all expansion regions into effective regions and invalid regions by comparing the effective probability with a preset value; wherein the effective probability represents the probability that the target expansion region is the region where the field edge is located; The detection module is used to increase the contrast of the grayscale inside each valid area through gamma transformation based on the effective probability of each valid area, and adaptively adjust the filter window to perform mean filtering on each invalid area. The size of the filter window is negatively correlated with the effective probability of the corresponding invalid area, so as to input the obtained target image into the pre-trained classification detection model and identify the bill type according to the output text area and content.
2. The fiscal and tax bill classification detection system based on image recognition according to claim 1 is characterized in that: The method for obtaining the effective probability of the target expansion area includes: Calculate the first index of the target expansion area : ; In the formula, is the area of the edge of the target expansion region; is the area of the target expansion region; is the median length of all edges in the target expansion area; is the width of the bounding rectangle of the target expansion area; is the information entropy of the target expansion area; The width of the circumscribed rectangle of the target expansion area is extended from both ends, and the newly added area is used as the expansion area. The second index of the target expansion area is calculated. The second index is positively correlated with the difference in information entropy between the target expansion area and the expansion area, and is negatively correlated with the width of the circumscribed rectangle and the degree of deviation from the overall level. The average of the first index and the second index is used as the effective probability of the target expansion area.
3. The fiscal and tax bill classification detection system based on image recognition according to claim 2 is characterized in that: The method for obtaining the second index of the target expansion area includes: Calculate the difference between the width of the bounding rectangle of the target expansion area and the median of the width of the bounding rectangles of all expansion areas, and normalize them to obtain the degree of deviation of the width of the bounding rectangle of the target expansion area from the overall level; The difference between 1 and the degree of deviation is calculated, and the difference and the difference in information entropy between the target expansion area and the expansion area are multiplied to obtain a second index of the target expansion area.
4. The fiscal and tax bill classification detection system based on image recognition according to claim 2 is characterized in that: The method for acquiring the information entropy of the target expansion area or the extended area comprises: The target expansion area or the extended area is divided into a plurality of square sub-areas with a set width, and the information entropy of the target expansion area or the extended area is calculated based on the probability that the number of edge pixels in all sub-areas takes different values and using the information entropy calculation formula.
5. The fiscal and tax bill classification detection system based on image recognition according to claim 1 is characterized in that: The step of increasing the contrast of the grayscale inside each effective area through gamma transformation based on the effective probability of each effective area includes: The sum of the effective probability of each effective area and 1 is used as a weight to weight the preset gamma coefficient to obtain an updated gamma coefficient of each effective area; The updated gamma coefficient is used to adjust the grayscale value of the pixel points in each effective area through gamma transformation.
6. The fiscal and tax bill classification detection system based on image recognition according to claim 1 is characterized in that: The adaptive adjustment of the filter window satisfies the following relationship: ; In the formula, For the The width of the adaptive filtering window of the invalid area; For the The effective probability of invalid regions; is the preset initial filter window width; The preset magnification factor; is the ceiling function.
7. The fiscal and tax bill classification detection system based on image recognition according to claim 1 is characterized in that: The method for obtaining the horizontal direction includes: Hough line detection is performed on the initial image, and based on the slope, all detected straight line edges whose length is greater than a preset length are clustered, so that the straight line direction corresponding to the cluster center of the category with a smaller average slope is used as the horizontal direction, and the straight line direction corresponding to the cluster center of another category is used as the vertical direction.
8. The fiscal and tax bill classification detection system based on image recognition according to claim 7 is characterized in that: The method for obtaining the non-straight edge comprises: Acquire all edges obtained by edge detection on the initial image, and all straight line edges obtained by Hough line detection on the initial image; All edges except the straight line edges among all the edges are defined as non-straight line edges.
9. The fiscal and tax bill classification detection system based on image recognition according to claim 7 is characterized in that: The length direction of the circumscribed rectangle of the target expansion area is parallel to the horizontal direction, and the width direction is parallel to the vertical direction.
10. The fiscal and tax bill classification detection system based on image recognition according to claim 1 is characterized in that: The classification detection model is an algorithm model that combines the target detection model of YOLOv10 with the OCR recognition model.
Citation Information
Patent Citations
Teaching-assistant book intelligent correcting and editing system based on character recognition
CN116071763A
Image segmentation method for edge correction loss based on topology key error identification
CN117593522A
Fiber diameter batch processing method
CN119984068A
Method for identifying character-strings
EP0479284A2