Automated Tax Monitoring and Warning Method and System
By performing completeness checks and edge profile identification on the tax warning documents, using the text database to fill in the loss targets, and generating monitoring warning tables, the problem of incomplete text in tax warning materials is solved, and efficient and accurate tax risk monitoring is achieved.
Patent Information
- Application Number
- CN202510484333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-17
AI Technical Summary
In the processing of existing tax warning materials, there is a situation where inefficient and difficult to accurately identify the incomplete text in the image file, resulting in inaccurate tax warning.
By performing integrity verification of tax warning files, identifying edge profiles and filling in loss targets, resuming file integrity using text databases, and generating monitoring warning tables to display information areas of different risk levels.
It improves the efficiency and accuracy of tax warning processing, ensures the integrity of document content, provides clear risk monitoring tools, and supports timely discovery and analysis of tax management.
Smart Images

Figure CN120013695B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to automatic processing technology, and in particular to an automated tax monitoring and early warning method and system. Background Art
[0002] In the prior art, for the processing of tax early warning materials, some adopt the method of manual review. This method not only has low efficiency but also is prone to omissions. Even if there are some automated processing means, there are still many deficiencies. For example, in terms of integrity verification, it is difficult for the prior art to accurately determine whether there is content missing in a file. Especially for a file containing image materials, it is impossible to effectively identify the incomplete situation of the text in the image, and the traditional processing method is difficult to meet the requirements of efficient and accurate processing.
[0003] Therefore, there is an urgent need for an automated tax monitoring and early warning method to improve processing efficiency and accuracy. Summary of the Invention
[0004] An embodiment of the present invention provides an automated tax monitoring and early warning method and system, which can provide an automated tax monitoring and early warning method to improve processing efficiency and accuracy.
[0005] In the first aspect of the embodiment of the present invention, an automated tax monitoring and early warning method is provided, including:
[0006] Perform integrity verification on all first tax early warning files in the tax early warning materials uploaded by the user in sequence. If it is determined that there is content missing, the first tax early warning file is taken as the second tax early warning file;
[0007] Perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform filling analysis on the loss target based on the text database to obtain the corresponding filling result;
[0008] Extract the text information in the first tax early warning file and the second tax early warning file respectively, and determine the tax type and type value corresponding to each text information;
[0009] Perform early warning monitoring processing based on the tax type and type value to generate a monitoring early warning table, and the monitoring early warning table at least includes information areas with different risk levels.
[0010] Optionally, in a possible implementation manner of the first aspect, the performing integrity verification on all first tax early warning files in the tax early warning materials uploaded by the user in sequence, and if it is determined that there is content missing, taking the first tax early warning file as the second tax early warning file includes:
[0011] If it is determined that there are image materials in the tax warning materials uploaded by the user, obtain the coordinates of all edge pixel points in the image materials to obtain an edge coordinate set;
[0012] Classify the edge coordinate set to obtain edge coordinate line subsets of 4 edge lines, generate extension lines corresponding to each edge line based on the edge coordinate line subsets, and use the intermediate area formed by the edge lines and the corresponding extension lines as the text mutilation recognition area;
[0013] If it is determined that there are mutilated characters in the text mutilation recognition area, it is determined that the integrity check fails and there is content missing, and the first tax warning file is used as the second tax warning file.
[0014] Optionally, in a possible implementation manner of the first aspect, the classifying the edge coordinate set to obtain edge coordinate line subsets of 4 edge lines, generating extension lines corresponding to each edge line based on the edge coordinate line subsets, and obtaining the text mutilation recognition area based on the edge lines and the corresponding extension lines includes:
[0015] Classify all pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain edge coordinate line subsets of 4 edge lines;
[0016] Determine the extension direction corresponding to each edge line, where each type of edge line has a preset extension direction;
[0017] Extend each pixel point in the edge line by a preset point position according to the extension direction to determine the corresponding extension points, and sequentially connect the adjacent extension points corresponding to each edge line to generate the corresponding extension lines.
[0018] Optionally, in a possible implementation manner of the first aspect, the classifying all pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain edge coordinate line subsets of 4 edge lines includes:
[0019] Determine the interval of the abscissas of all pixel points, determine the maximum ordinate value of all pixel points in the interval of the abscissas to obtain the edge coordinate line subset of the first edge, and determine the minimum ordinate value of all pixel points in the interval of the abscissas to obtain the edge coordinate line subset of the second edge;
[0020] Determine the interval of the ordinates of all pixel points, determine the maximum abscissa value of all pixel points in the interval of the ordinates to obtain the edge coordinate line subset of the third edge, and determine the minimum abscissa value of all pixel points in the interval of the ordinates to obtain the edge coordinate line subset of the fourth edge.
[0021] Optionally, in a possible implementation manner of the first aspect, if it is determined that there are incomplete characters in the incomplete character recognition area, and it is determined that there is content missing due to the failure of the integrity check, using the first tax warning file as the second tax warning file includes:
[0022] Obtain the number of pixel points in the incomplete character recognition area. If it is determined that the number of pixel points is empty, it is determined that the integrity check is passed;
[0023] If it is determined that the number of pixel points is not empty, obtain the coordinate values of the abscissa and ordinate of all pixel points. If the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, it is determined that the integrity check is passed;
[0024] If the coordinate values of the abscissa and ordinate of all pixel points do not pass the regularity check, it is determined that there are incomplete characters in the incomplete character recognition area, and it is determined that the integrity check is not passed.
[0025] Optionally, in a possible implementation manner of the first aspect, the step of obtaining the coordinate values of the abscissa and ordinate of all pixel points, and if the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, determining that the integrity check is passed includes:
[0026] Obtain all pixel points with a preset pixel value as the points to be verified;
[0027] If it is determined that the ordinate values of the points to be verified with adjacent abscissas are the same, or the absolute value of the difference in ordinate values is the same preset value, it is determined that the regularity check is passed;
[0028] If it is determined that the abscissa values of the points to be verified with adjacent ordinates are the same, or the absolute value of the difference in abscissa values is the same preset value, it is determined that the regularity check is passed;
[0029] If it is determined that the points to be verified are less than the preset value, it is determined that the regularity check is passed.
[0030] Optionally, in a possible implementation manner of the first aspect, performing edge contour recognition on the second tax warning file, determining the loss target superimposed on the edge contour, and obtaining the corresponding filling result through filling analysis of the loss target based on the character database includes:
[0031] Obtain the characters in the second tax warning file that are not edge contours, and select the center point of a non-edge character in each row as the alignment point;
[0032] Generate a standard character comparison box preset for the non-edge characters. If there are multiple standard character comparison boxes, align the center points of the standard character comparison boxes with the alignment points in sequence, and determine the first edge abscissas of the characters on both sides corresponding to the alignment point in the direction close to the alignment point. There are 2 first edge abscissas;
[0033] Obtaining the second edge horizontal coordinates of both sides of the standard text comparison frame, the second edge horizontal coordinates being 2, and obtaining the third edge horizontal coordinates of both sides of the standard text, the third edge horizontal coordinates being 2;
[0034] Based on the numerical relationship between the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate, the standard text comparison frame is screened to obtain the target text comparison frame, and based on the target text comparison frame, the lost target is superimposed and then a completion analysis is performed to obtain the corresponding completion result.
[0035] Optionally, in a possible implementation manner of the first aspect, the method of screening the standard text comparison frame based on the numerical relationship among the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result includes:
[0036] Taking the center point as the dividing point, the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate are divided toward both sides to obtain two sets of edge horizontal coordinate sets;
[0037] Calculate the difference between the first edge horizontal coordinate and the second edge horizontal coordinate in each set of edge horizontal coordinates to obtain a first difference, and calculate the difference between the second edge horizontal coordinate and the third edge horizontal coordinate to obtain a second difference;
[0038] The standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship is determined to obtain the target text comparison frame, and the loss target is superimposed based on the target text comparison frame and then a completion analysis is performed to obtain the corresponding completion result.
[0039] Optionally, in a possible implementation manner of the first aspect, determining a standard text comparison frame corresponding to the first difference and the second difference having the closest absolute value relationship to obtain a target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result, including:
[0040] Determine the properties of the incomplete text recognition area where the loss target is located to obtain the alignment of the target text comparison frame, and the properties of each incomplete text recognition area have a preset alignment;
[0041] Based on the preset alignment method, the target text comparison frame and the loss target are superimposed to obtain the position and shape of the loss target in the target text comparison frame;
[0042] Based on the position and glyph of the loss target in the target text comparison frame, the text database is traversed to perform text comparison analysis, the text corresponding to the position and glyph in the target text comparison frame is determined as the completion text, and the completion result is obtained.
[0043] Optionally, in a possible implementation of the first aspect, the method further includes:
[0044] The step of traversing the text database to perform text comparison analysis based on the position and glyph of the lost target in the target text comparison frame, determining the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtaining the completion result includes:
[0045] Based on the target text comparison frame, the incomplete text recognition area is extended once, and the completed text is filled into the target text comparison frame;
[0046] If there are multiple complementing characters, a secondary extension process is performed on the corresponding target character comparison frame to obtain a new target character comparison frame and all the complementing characters are filled in sequentially, so that each loss target has a corresponding complementing character;
[0047] If the completion text of any loss target is empty, fill in the blank in the corresponding target text comparison box.
[0048] Optionally, in a possible implementation manner of the first aspect, the monitoring and early warning table is generated after the early warning monitoring processing is performed based on the tax type and the type value, and the monitoring and early warning table at least includes information areas of different risk levels, including:
[0049] Initialize a monitoring and warning table, wherein the monitoring and warning table includes low-risk areas and high-risk areas;
[0050] Count all tax types and type values that meet the requirements and fill them into the low-risk area, and set the corresponding contents in the first tax warning file and the second tax warning file to correspond to the information in the low-risk area;
[0051] All tax types and type values that do not meet the requirements are counted and filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set corresponding to the information in the high-risk area.
[0052] Optionally, in a possible implementation of the first aspect, the counting of all tax types and type values that do not meet the requirements and filling them into the high-risk area, and setting the corresponding contents in the first tax early warning file and the second tax early warning file to correspond to the information of the high-risk area, includes:
[0053] Count all tax types and type values that do not meet the requirements and fill them into the high-risk area;
[0054] If the characters of the tax types and type values that do not meet the judgment exist in the target text comparison box, add a completion verification label to the corresponding tax types and type values.
[0055] In the second aspect of the embodiments of the present invention, an automated tax monitoring and early warning system is provided, including:
[0056] A verification module for sequentially performing integrity verification on all the first tax early warning files in the tax early warning materials uploaded by the user. If it is determined that there is content missing, the first tax early warning file is used as the second tax early warning file;
[0057] An identification module for performing edge contour identification on the second tax early warning file, determining the loss target superimposed on the edge contour, and obtaining the corresponding completion result by performing completion analysis on the loss target based on the text database;
[0058] An extraction module for extracting the text information in the first tax early warning file and the second tax early warning file respectively, and determining the tax type and type value corresponding to each piece of text information;
[0059] An early warning module for generating a monitoring and early warning table after performing early warning monitoring processing based on the tax type and type value. The monitoring and early warning table at least includes information areas with different risk levels. Beneficial effects
[0060] By performing integrity verification on the first tax early warning file in the tax early warning materials uploaded by the user, the present invention can accurately determine whether there is content missing in the file. For files containing image materials, by obtaining the coordinates of all edge pixel points in the image materials, classifying the edge coordinate set, determining the text mutilation recognition area, and further determining whether there are mutilated characters in this area, accurate verification of the file integrity is achieved. Through the analysis and processing of the coordinates, the text mutilation area is located and it is determined whether the text is complete. If it is found that there are mutilated characters in the text mutilation recognition area, the file is marked as the second tax early warning file for subsequent targeted processing, effectively avoiding the problem of inaccurate tax early warning caused by missing file content and improving the quality of tax early warning data.
[0061] After the first tax warning file is determined to have missing content and marked as the second tax warning file, the present invention can effectively restore the integrity of the file by performing edge contour recognition on it, determining the loss targets superimposed on the edge contours, and conducting filling analysis on the loss targets based on the text database. Specifically, by obtaining the text within the second tax warning file that is not part of the edge contour, determining the alignment points, generating a standard text comparison box, and screening the target text comparison boxes, the text filling of the loss targets is ultimately achieved, making the text information in the file complete and accurate, providing a reliable data basis for subsequent tax analysis.
[0062] The present invention extracts the text information within the first tax warning file and the second tax warning file respectively, determines the tax types and type values corresponding to each piece of text information, and conducts warning monitoring and processing based on this information to generate a monitoring and warning table containing information regions with different risk levels. The system can accurately count the tax types and values that meet the requirements and those that do not, fill them into the low-risk region and the high-risk region respectively, and set the corresponding content in the file to correspond to the region information. At the same time, for tax information that does not meet the requirements and the characters of which are within the target text comparison box, a filling verification label is added, further improving the accuracy of warning monitoring. The monitoring and warning table can clearly display the risk levels of different tax types and values. Tax administrators can intuitively understand the tax risk situation, promptly discover high-risk tax information, and trace and analyze its source and specific circumstances, providing strong support for tax decision-making and effectively improving the efficiency and accuracy of tax management. Brief Description of the Drawings
[0063] Figure 1 is a flowchart of an automated tax monitoring and warning method provided by an embodiment of the present invention
[0064] Figure 2 is a schematic diagram of an extension line provided by an embodiment of the present invention;
[0065] Figure 3 is a schematic diagram of the structure of an automated tax monitoring and warning system provided by an embodiment of the present invention Detailed Embodiment
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] See Figure 1, is a flow chart of an automated tax monitoring and early warning method provided by an embodiment of the present invention, the method comprising:
[0068] S1, perform integrity checks on all first tax warning files in the tax warning materials uploaded by the user in turn. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file.
[0069] In the automated tax monitoring and early warning system, the integrity check of the first tax warning file in the tax warning materials uploaded by the user is required to ensure the quality of the tax warning data and the accuracy of subsequent analysis and early warning. When it is found that the file has missing content, the file will be marked as the second tax warning file for subsequent targeted processing.
[0070] In some embodiments, the integrity check of all first tax warning files in the tax warning materials uploaded by the user is performed in sequence, and if it is determined that there is content missing, the first tax warning file is used as the second tax warning file, including:
[0071] S11, if it is determined that the tax warning materials uploaded by the user contain image materials, the coordinates of all edge pixel points in the image materials are obtained to obtain an edge coordinate set.
[0072] The image material can be an example of tax materials such as invoice materials. Edge refers to the edge of the image. When the system determines that the tax warning materials uploaded by the user contain image materials, it will obtain the coordinates of the edge pixels. Scan the image material and accurately extract the coordinates of all edge pixels in the image to form an edge coordinate set. The edge coordinate set is required for subsequent analysis of the image structure and recognition of the incomplete text area.
[0073] S12, classifying the edge coordinate set to obtain edge coordinate line subsets of four edge lines, generating an extension line corresponding to each edge line based on the edge coordinate line subsets, and using the middle area formed by the edge line and the corresponding extension line as the text incomplete recognition area.
[0074] See also Figure 2 The system will classify the acquired edge coordinate set to form 4 edge coordinate line subsets. Based on the subsets, the extension line corresponding to each edge line is generated. The edge line and its corresponding extension line are used to determine the text defect recognition area, providing a prerequisite for the subsequent accurate positioning of the text defect position.
[0075] The step of classifying the edge coordinate set to obtain edge coordinate line subsets of four edge lines, generating an extension line corresponding to each edge line based on the edge coordinate line subsets, and obtaining a text incomplete recognition area based on the edge line and the corresponding extension line includes:
[0076] S121. Classify all the pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain four subsets of edge coordinate lines of the four edge lines.
[0077] To form four subsets of edge coordinate lines of the four edge lines, the system classifies all the pixel points in the edge coordinate set according to the abscissa value and the ordinate value. By sorting out the coordinate data, the pixel points are classified into the corresponding subsets of edge line coordinates, clearly defining different parts of the image edge.
[0078] Among them, the step of classifying all the pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain four subsets of edge coordinate lines of the four edge lines includes:
[0079] S1211. Determine the interval of the abscissa of all the pixel points, determine the maximum ordinate value of all the pixel points within the interval of the abscissa to obtain the subset of edge coordinate lines of the first edge, and determine the minimum ordinate value of all the pixel points within the interval of the abscissa to obtain the subset of edge coordinate lines of the second edge.
[0080] The system first determines the interval of the abscissa of all the pixel points. Within this interval, find the maximum ordinate value of the pixel points within the abscissa interval, and classify these pixel points into the subset of edge coordinate lines of the first edge, that is, the topmost line; find the minimum ordinate value of the pixel points within the abscissa interval and classify them into the subset of edge coordinate lines of the second edge, that is, the bottommost line, realizing a preliminary division of the image edge from the vertical dimension.
[0081] S1212. Determine the interval of the ordinate of all the pixel points, determine the maximum abscissa value of all the pixel points within the interval of the ordinate to obtain the subset of edge coordinate lines of the third edge, and determine the minimum abscissa value of all the pixel points within the interval of the ordinate to obtain the subset of edge coordinate lines of the fourth edge.
[0082] The system determines the interval of the ordinate of all the pixel points. Within this interval, find the maximum abscissa value of the pixel points within the ordinate interval and classify them into the subset of edge coordinate lines of the third edge, that is, the rightmost line; find the minimum abscissa value of the pixel points within the ordinate interval and classify them into the subset of edge coordinate lines of the fourth edge, that is, the leftmost line, further refining the division of the image edge from the horizontal dimension.
[0083] S122. Determine the extension direction corresponding to each edge line, where each type of edge line has a preset extension direction.
[0084] Among them, the extension direction is preset. For example, the topmost edge line extends downward, the bottommost line extends upward, the leftmost edge line extends to the right, and the rightmost line extends to the left.
[0085] S123. For each pixel point within the edge line, extend it by a preset number of points in the extension direction to determine the corresponding extended points, and sequentially connect the adjacent extended points corresponding to each edge line to generate the corresponding extended line.
[0086] The system, according to the determined extension direction, moves each pixel point within the edge line along the extension direction by a preset number of points to determine the corresponding extended points. Sequentially connect the adjacent extended points corresponding to each edge line to generate the corresponding extended line, thereby constructing the text mutilation recognition area to ensure accurate recognition of text mutilation in the follow-up.
[0087] S13. If it is determined that there are mutilated characters in the text mutilation recognition area, then it is judged that the integrity check fails and there is content missing, and the first tax warning file is used as the second tax warning file.
[0088] The system determines whether there are mutilated characters in this area to further determine whether the file passes the integrity check. If it is judged that there are mutilated characters, it means that the file fails the integrity check and there is content missing. Furthermore, mark the first tax warning file as the second tax warning file for subsequent special processing; if there are no mutilated characters, it is considered that the file passes the integrity check and can enter the subsequent normal tax warning analysis process.
[0089] Among them, the statement that if it is determined that there are mutilated characters in the text mutilation recognition area, then it is judged that the integrity check fails and there is content missing, and the first tax warning file is used as the second tax warning file includes:
[0090] S131. Obtain the number of pixel points within the text mutilation recognition area. If it is judged that the number of pixel points is empty, then it is judged that the integrity check is passed.
[0091] The system first obtains the number of pixel points within the text mutilation recognition area. If the number of pixel points in this area is empty, that is, there are no pixel points, which means that no text-related information is detected in this area. It can be initially judged that there is no text mutilation in this area, and thus it is judged that the first tax warning file passes the integrity check. Because under normal circumstances, if the text is complete, a certain number of pixel points will not be detected in the recognition area to form text. If the number of pixel points is empty, it is very likely that the text is complete and there is no mutilation.
[0092] S132. If it is judged that the number of pixel points is not empty, then obtain the coordinate values of the abscissa and ordinate of all pixel points. If the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, then it is judged that the integrity check is passed.
[0093] When it is determined that the number of pixel points in the text mutilation recognition area is not empty, the system further obtains the coordinate values of the abscissa and ordinate of all pixel points, and performs a regularity check on these coordinate values to more accurately determine whether the text is mutilated. If the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, it is determined that the first tax warning file passes the integrity check.
[0094] Among them, the step of obtaining the coordinate values of the abscissa and ordinate of all pixel points, and if the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, it is determined that the integrity check is passed, includes:
[0095] S1321, obtain all pixel points with a preset pixel value as the points to be verified.
[0096] The system first obtains all pixel points with a preset pixel value as the points to be verified. The preset pixel value is set according to the display characteristics of the text in the image and past experience, etc. These pixel points with specific pixel values are more likely to be the key pixel points that make up the text. For example, it is the pixel value corresponding to black.
[0097] S1322, if it is determined that the ordinate values of the points to be verified with adjacent abscissas are the same, or the absolute value of the difference in ordinate values is the same preset value, it is determined that the regularity check is passed.
[0098] This step is described for the text mutilation recognition areas above and below.
[0099] The system judges the ordinate values of the points to be verified with adjacent abscissas. If it is determined that the ordinate values of the points to be verified with adjacent abscissas are the same, for example, in the form of a horizontal line, or the absolute value of the difference in ordinate values is the same preset value, for example, in the form of a broken line.
[0100] The above indicates that these points to be verified show a certain regularity in the vertical direction, and it can be initially judged that it is not in the form of text, but in the form of a horizontal line or a broken line that often appears in some documents. For example, invoice documents often have wireframes.
[0101] S1323, if it is determined that the abscissa values of the points to be verified with adjacent ordinates are the same, or the absolute value of the difference in abscissa values is the same preset value, it is determined that the regularity check is passed;
[0102] This step is described for the text mutilation recognition areas on the left and right.
[0103] The system judges the abscissa values of the points to be verified with adjacent ordinates. If it is determined that the abscissa values of the points to be verified with adjacent ordinates are the same, for example, in the form of a vertical line, or the absolute value of the difference in abscissa values is the same preset value, for example, in the form of a broken line.
[0104] The above shows that these points to be verified show a certain regularity in the vertical direction. It can be preliminarily judged that they are not in the form of text, but in the form of horizontal lines or broken lines that are often found in some documents. For example, invoice documents often have wireframes.
[0105] S1324, if it is determined that the point to be verified is less than the preset value, then the regular verification is passed.
[0106] The system determines whether the number of points to be verified is less than the preset value. If it is determined that the point to be verified is less than the preset value, it means that the number of key pixel points constituting the text in this area is small, and the regular verification is passed.
[0107] S133, if there are coordinate values of the abscissa and ordinate of the pixel points that do not pass the regular verification, it is determined that there are incomplete characters in the text incomplete recognition area, and it is determined that the integrity verification is not passed.
[0108] If there is a situation where the coordinate values do not pass the above regular verification during the process of regular verification of the coordinate values of the abscissa and ordinate of the pixel points, the system determines that there are incomplete characters in the text incomplete recognition area, and further determines that the first tax warning file does not pass the integrity verification. This means that the text of the image material in the file is incomplete, which may affect the accuracy of tax warning analysis. Therefore, the file is marked as the second tax warning file.
[0109] S2, perform edge contour recognition on the second tax warning file, determine the loss targets superimposed on the edge contours, and perform filling analysis on the loss targets based on the text database to obtain the corresponding filling results.
[0110] After the first tax warning file is determined to have content missing and marked as the second tax warning file, the system will perform a further processing flow on it. First, perform edge contour recognition on the second tax warning file to determine the edge contours of the image part. On this basis, determine the loss targets superimposed on these edge contours. These loss targets are usually the missing parts of information caused by incomplete characters. Subsequently, the system performs filling analysis on the loss targets based on the pre-established text database. The text database stores the characteristics and information of various standard texts. The system matches and analyzes the loss targets with the texts in the database to obtain the corresponding filling results, thereby restoring the missing text information in the file and improving the integrity and accuracy of the tax warning file.
[0111] In some embodiments, the performing edge contour recognition on the second tax warning file, determining the loss targets superimposed on the edge contours, and performing filling analysis on the loss targets based on the text database to obtain the corresponding filling results includes:
[0112] S21. Obtain the text of non-edge contours in the second tax warning file, and select the center point of a non-edge text in each line as the alignment point.
[0113] The system extracts the text content of non-edge contours from the second tax warning file. For each line of non-edge text, the system selects the center point of one text as the alignment point. By determining the alignment point, it provides a reference point for generating a standard text comparison box later, enabling more accurate subsequent text analysis and processing, and helping to improve the accuracy of the filling analysis.
[0114] Among them, the way to obtain the center point of the text can be to determine its leftmost point coordinate and rightmost point coordinate, determine the abscissa value through the leftmost point coordinate and rightmost point coordinate, then determine its topmost point coordinate and bottommost point coordinate, and determine the ordinate value through the topmost point coordinate and bottommost point coordinate, and combine the abscissa value and ordinate value to obtain the center point.
[0115] S22. Generate a standard text comparison box preset for non-edge text. If there are multiple standard text comparison boxes, align the center points of the standard text comparison boxes with the alignment point in sequence, and determine the first edge abscissas of the text on both sides of the alignment point close to the alignment point direction, and there are 2 first edge abscissas.
[0116] The system generates a standard text comparison box preset for non-edge text. If multiple standard text comparison boxes are generated, the system will align the center points of each standard text comparison box with the previously determined alignment point in sequence. After alignment, the system determines the first edge abscissas of the text on both sides of the alignment point close to the alignment point direction. One first edge abscissa is obtained in each direction, and there are two in total. These first edge abscissas reflect the boundary positions of the text related to the alignment point in the horizontal direction.
[0117] S23. Obtain the second edge abscissas on both sides of the standard text comparison box, and there are 2 second edge abscissas, and obtain the third edge abscissas on both sides of the standard text, and there are 2 third edge abscissas.
[0118] The system further obtains the second edge abscissas on both sides of the standard text comparison box, also one in each direction, and there are two in total. These two second edge abscissas determine the boundary range of the standard text comparison box itself in the horizontal direction. At the same time, the system obtains the third edge abscissas on both sides of the standard text, also one in each direction, and there are two in total. The third edge abscissas provide more detailed position information about the standard text in the horizontal direction. By obtaining these different levels of edge abscissas, the system can comprehensively understand the spatial positions and ranges of the standard text comparison box and the standard text in the horizontal direction.
[0119] S24. Based on the numerical relationship among the first edge abscissa, the second edge abscissa, and the third edge abscissa, filter the standard text comparison box to obtain the target text comparison box, and perform complement analysis on the superimposed loss target based on the target text comparison box to obtain the corresponding complement result.
[0120] After obtaining the first edge abscissa, the second edge abscissa, and the third edge abscissa, the system filters the standard text comparison box based on the numerical relationship among these abscissas to determine the most suitable target text comparison box. After determining the target text comparison box, superimpose it with the loss target, and perform a detailed complement analysis on the loss target based on the text database to obtain the corresponding complement result and restore the missing text information in the second tax warning file.
[0121] Among them, the step of filtering the standard text comparison box based on the numerical relationship among the first edge abscissa, the second edge abscissa, and the third edge abscissa to obtain the target text comparison box, and performing complement analysis on the superimposed loss target based on the target text comparison box to obtain the corresponding complement result includes:
[0122] S241. Using the center point as the division point, divide the first edge abscissa, the second edge abscissa, and the third edge abscissa to both sides to obtain two sets of edge abscissa sets.
[0123] The system uses the previously determined center point (such as the alignment point) as the division point and divides the first edge abscissa, the second edge abscissa, and the third edge abscissa to both sides. Through this division method, these abscissas are divided into two sets of edge abscissa sets. This grouping operation helps to analyze and compare the abscissas on both sides subsequently, making the screening of the standard text comparison box more meticulous and accurate.
[0124] S242. Calculate the difference between the first edge abscissa and the second edge abscissa within each set of edge abscissa sets to obtain the first difference, and calculate the difference between the second edge abscissa and the third edge abscissa to obtain the second difference.
[0125] For each set of edge abscissa sets, the system calculates the difference between the first edge abscissa and the second edge abscissa within the set to obtain the first difference. At the same time, calculate the difference between the second edge abscissa and the third edge abscissa to obtain the second difference. These differences reflect the distance relationship between different edge abscissas. By calculating these differences, the characteristics of the standard text comparison box in the horizontal direction can be quantified. The calculation results of these differences will be an important basis for subsequent screening of the target text comparison box, helping the system determine which standard text comparison box best matches the actual text layout in the file.
[0126] S243, determining the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result.
[0127] The system compares the absolute value relationship between the first difference and the second difference corresponding to all standard text comparison frames, finds the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship, and determines it as the target text comparison frame. This target text comparison frame is considered to be the most consistent with the actual text layout and loss target in the file. After determining the target text comparison frame, the system superimposes the loss target based on the target text comparison frame and performs a completion analysis to obtain the corresponding completion result.
[0128] The step of determining the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain the corresponding completion result includes:
[0129] S2431, determining the attributes of the incomplete text recognition region where the loss target is located to obtain the alignment of the target text comparison frame, and the attributes of each incomplete text recognition region have a preset alignment.
[0130] The system first determines the properties of the text defect recognition area where the loss target is located. The properties of each text defect recognition area are pre-set with the corresponding alignment. By determining the properties of the area where the loss target is located, the system can obtain the alignment of the target text comparison box. This alignment determination helps to accurately superimpose the target text comparison box with the loss target in the future, ensuring the accuracy of the completion analysis.
[0131] The attribute of each incomplete text recognition area having a preset alignment means that when the incomplete text is located in different areas, the incomplete position may be different. For example, for the incomplete text in the left incomplete text recognition area, the missing part is generally the left part of the text; for the incomplete text in the right incomplete text recognition area, the missing part is generally the right part of the text, and so on.
[0132] S2432, based on the preset alignment method, the target text comparison frame and the loss target are superimposed to obtain the position and font shape of the loss target in the target text comparison frame.
[0133] Based on the preset alignment method, the system overlays the target text comparison frame with the loss target. The system obtains the specific position information of the loss target in the target text comparison frame, as well as the glyph features presented by the loss target. These position and glyph information are key data for subsequent text comparison analysis, which can help the system accurately find text that matches the loss target in the text database.
[0134] S2433, based on the position and glyph of the lost target in the target text comparison frame, traverse the text database to perform text comparison analysis, determine the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtain the completion result.
[0135] After obtaining the position and glyph information of the loss target in the target text comparison box, the system traverses the text database based on this information and performs detailed text comparison analysis. The purpose is to find the text that matches the position and glyph in the target text comparison box, and determine it as the completion text. Finally, accurate completion results are obtained to complete the text completion work of the loss target in the second tax warning file.
[0136] The step of traversing the text database to perform text comparison analysis based on the position and glyph of the lost target in the target text comparison frame, determining the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtaining the completion result includes:
[0137] S24331, performing an extension process on the incomplete text recognition area based on the target text comparison frame, and filling the completed text into the target text comparison frame;
[0138] The system first performs an extension process on the incomplete text recognition area based on the target text comparison frame. This extension process is to expand the scope of the target text comparison frame so that it can better accommodate possible complementary text. After completing the extension process, the system fills the determined complementary text into the target text comparison frame. In this way, the text complement of the lost target is initially achieved, making the text information in the target text comparison frame more complete.
[0139] S24332, if there are multiple complementing characters, a secondary extension process is performed on the corresponding target character comparison frame to obtain a new target character comparison frame and all the complementing characters are filled in sequentially, so that each loss target has a corresponding complementing character;
[0140] When there are multiple complementary characters, it means that the loss target may have multiple matching characters. At this time, the system performs secondary extension processing on the corresponding target character comparison box to generate a new target character comparison box. The secondary extension processing is to further expand the scope of the target character comparison box to meet the filling requirements of multiple complementary characters.
[0141] For S24333, if the supplementary text for any loss target is empty, fill in "empty" in the corresponding target text comparison box.
[0142] After traversing the text database for text comparison and analysis, if it is found that the supplementary text for any loss target is empty, that is, no text corresponding to the position and glyph of the loss target is found, the system will fill in "empty" in the corresponding target text comparison box. This processing method ensures data consistency and accuracy, clearly identifies the situation where the loss target cannot be supplemented through the existing text database, and provides clear information for subsequent processing. For example, if there is no matching text in the text database, the system will fill in "empty" in the corresponding target text comparison box so that the staff can discover it in time and take further measures, such as manual review or supplementing the text database.
[0143] S3. Extract the text information in the first tax warning file and the second tax warning file respectively, and determine the tax type and type value corresponding to each text information.
[0144] The system processes the first tax warning file and the second tax warning file respectively, extracts the text information in the file through text recognition technology. After extracting the text information, the system further analyzes each text information, and determines its corresponding tax type and type value according to the pre-set rules and standards. For example, for text information related to the amount, the system will judge which tax item (such as value-added tax, income tax, etc.) it belongs to and the specific amount value. Through this step, the system converts the text information in the tax file into structured data that can be used for warning monitoring, providing basic data support for the subsequent generation of the monitoring warning table.
[0145] S4. After performing warning monitoring processing based on the tax type and type value, generate a monitoring warning table, and the monitoring warning table at least includes information areas with different risk levels.
[0146] After obtaining the tax type and type value, the system performs warning monitoring processing based on this information and finally generates a monitoring warning table. The monitoring warning table is an important output of tax warning analysis, and it at least includes information areas with different risk levels so that users can intuitively understand the tax risk situation.
[0147] In some embodiments, the step of generating a monitoring warning table after performing warning monitoring processing based on the tax type and type value, and the monitoring warning table at least includes information areas with different risk levels, includes:
[0148] S41. Initialize the monitoring warning table, and the monitoring warning table includes a low-risk area and a high-risk area.
[0149] The system first initializes the monitoring and warning table. During the initialization process, the monitoring and warning table is set to include at least two basic information areas: the low-risk area and the high-risk area. The low-risk area is used to display tax information that meets the requirements and has relatively low risks; the high-risk area is used to present tax information that does not meet the requirements and may have relatively high tax risks. This way of area division provides a framework for the subsequent classification and display of tax information.
[0150] S42. Statistically count all tax types and type values that meet the requirements and fill them into the low-risk area, and correspondingly set the corresponding contents in the first tax warning file and the second tax warning file with the information in the low-risk area.
[0151] The system statistically counts all tax types and type values, and filters out tax information that meets the requirements. Here, "meeting the requirements" is judged according to pre-set criteria, such as the compliance of tax types, whether the type values are within the normal range, etc. Fill the filtered tax types and type values that meet the requirements into the low-risk area of the monitoring and warning table. At the same time, the system correspondingly sets the corresponding contents in the first tax warning file and the second tax warning file with the information in the low-risk area, so that users can clearly see the sources and locations of these low-risk tax information in the original files. In this way, users can quickly understand which tax information is in a low-risk state and the specific situations of the relevant information in the files.
[0152] S43. Statistically count all tax types and type values that do not meet the requirements and fill them into the high-risk area, and correspondingly set the corresponding contents in the first tax warning file and the second tax warning file with the information in the high-risk area.
[0153] The system statistically counts tax types and type values again, and this time filters out tax information that does not meet the requirements. These situations that do not meet the requirements may include incorrect tax types, abnormal type values, etc. Fill the filtered tax types and type values that do not meet the requirements into the high-risk area of the monitoring and warning table. Similarly, the system correspondingly sets the corresponding contents in the first tax warning file and the second tax warning file with the information in the high-risk area, so that users can accurately trace the sources and specific situations of high-risk tax information. In this way, users can quickly pay attention to the parts that may have tax risks and further analyze and process these high-risk information. Finally, the monitoring and warning table generated through the above steps provides a visual and clear tax risk monitoring tool for tax management personnel, which helps to discover and solve tax problems in a timely manner.
[0154] Among them, all tax types and type values that do not meet the requirements are counted and filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set corresponding to the information in the high-risk area, including:
[0155] S431, count all tax types and type values that do not meet the requirements and fill them into the high-risk area.
[0156] The system comprehensively counts all tax types and type values, and judges which tax types and values do not meet the requirements according to the preset criteria. After screening out all tax types and type values determined to be non-compliant, the system fills them into the high-risk area of the monitoring and warning form. In this way, the high-risk area centrally displays all information that may have tax risks, enabling tax administrators to quickly locate and focus on these key data.
[0157] S432, if the characters of the tax types and type values judged to be non-compliant exist in the target text comparison box, add a completion verification label to the corresponding tax types and type values.
[0158] After filling the tax types and type values that do not meet the requirements into the high-risk area, the system further checks this information. For each tax type and type value that does not meet the requirements, the system judges whether its characters exist in the target text comparison box generated during the previous processing. The target text comparison box is generated during the completion analysis of the second tax warning file and is used to determine and process incomplete or missing parts of the text. If it is judged that the characters of the tax types and type values that do not meet the requirements exist in the target text comparison box, this means that the tax information may not meet the requirements due to incomplete or missing text. To further verify and process this situation, the system adds a completion verification label to the corresponding tax types and type values. After adding the completion verification label, tax administrators can more clearly identify which high-risk tax information may be related to text processing, and thus specifically recheck this information to see if there are inaccuracies or incompleteness in text completion, so as to ensure the accuracy and integrity of tax information and further reduce tax risks.
[0159] See Figure 3 , which is a schematic structural diagram of an automated tax monitoring and warning system provided by an embodiment of the present invention. The system includes:
[0160] A verification module for sequentially performing integrity verification on all the first tax warning files in the tax warning materials uploaded by the user. If it is judged that there is content missing, the first tax warning file will be used as the second tax warning file;
[0161] An identification module, configured to perform edge contour identification on the second tax warning file, determine a loss target superimposed on the edge contour, and perform complement analysis on the loss target based on a character database to obtain a corresponding complement result;
[0162] An extraction module, configured to extract text information in the first tax warning file and the second tax warning file respectively, and determine the tax type and type value corresponding to each piece of text information;
[0163] A warning module, configured to generate a monitoring warning table after performing warning monitoring processing based on the tax type and type value, where the monitoring warning table at least includes information areas with different risk levels.
[0164] The present invention also provides a storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.
[0165] Among them, the storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transmission of a computer program from one place to another. A computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also exist as discrete components in a communication device. The storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0166] The present invention also provides a program product, which includes execution instructions stored in a storage medium. At least one processor of the device can read the execution instructions from the storage medium, and at least one processor executes the execution instructions to enable the device to implement the methods provided by the above various embodiments.
[0167] In the above embodiments of the terminal or the server, it should be understood that the processor may be a central processing unit (Central Processing Unit, CPU for short), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP for short), application specific integrated circuits (Application Specific Integrated Circuit, ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the present invention can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An automated tax monitoring and warning method, characterized in that, Including: Successively perform integrity verification on all the first tax warning documents in the tax warning materials uploaded by the user. If it is determined that there is missing content, then regard the first tax warning document as the second tax warning document; Perform edge contour recognition on the second tax warning document, determine the loss target superimposed on the edge contour, and perform filling analysis on the loss target based on the character database to obtain the corresponding filling result, including: Obtain the characters in the second tax warning document that are not the edge contour, and select the center point of one non-edge character in each line as the alignment point; Generate a standard character comparison box preset for the non-edge characters. If there are multiple standard character comparison boxes, then successively align the center points of the standard character comparison boxes with the alignment points, and determine the first edge abscissas of the characters on both sides of the alignment point close to the alignment point direction. There are 2 first edge abscissas; Obtain the second edge abscissas on both sides of the standard character comparison box. There are 2 second edge abscissas, and obtain the third edge abscissas on both sides of the standard characters. There are 2 third edge abscissas; Based on the numerical relationship among the first edge abscissa, the second edge abscissa, and the third edge abscissa, screen the standard character comparison box to obtain the target character comparison box. Based on the target character comparison box, perform filling analysis on the superimposed loss target to obtain the corresponding filling result, including: Taking the center point as the division point, divide the first edge abscissa, the second edge abscissa, and the third edge abscissa to both sides to obtain two sets of edge abscissa sets; Calculate the difference between the first edge abscissa and the second edge abscissa in each set of edge abscissa sets to obtain the first difference, and calculate the difference between the second edge abscissa and the third edge abscissa to obtain the second difference; Determine the standard character comparison box corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target character comparison box. Based on the target character comparison box, perform filling analysis on the superimposed loss target to obtain the corresponding filling result, including: Based on the preset alignment method, perform superposition processing on the target character comparison box and the loss target, and obtain the position and glyph of the loss target in the target character comparison box; Based on the position and glyph of the loss target in the target character comparison box, traverse the character database for character comparison analysis, determine the character corresponding to the position and glyph in the target character comparison box as the filling character, and obtain the filling result; Extract the character information in the first tax warning document and the second tax warning document respectively, and determine the tax type and type value corresponding to each piece of character information; Perform warning monitoring processing based on the tax type and type value to generate a monitoring warning table, and the monitoring warning table at least includes information areas with different risk levels.
2. The automated tax monitoring and warning method according to claim 1, wherein The step of successively performing integrity verification on all the first tax warning documents in the tax warning materials uploaded by the user, and if it is determined that there is missing content, then regarding the first tax warning document as the second tax warning document, includes: If it is determined that there are image materials in the tax warning materials uploaded by the user, obtain the coordinates of all edge pixel points in the image materials to obtain an edge coordinate set; Classify the edge coordinate set to obtain four subsets of edge coordinate lines for the four edge lines. Generate an extension line corresponding to each edge line based on the subset of edge coordinate lines. Use the intermediate area formed by the edge line and the corresponding extension line as the text mutilation recognition area, including Determine the extension direction corresponding to each edge line, where each type of edge line has a preset extension direction; Extend each pixel point in the edge line by a preset number of positions in the extension direction to determine the corresponding extension points. Connect the adjacent extension points corresponding to each edge line in sequence to generate the corresponding extension line; If it is determined that there are mutilated characters in the text mutilation recognition area, it is determined that the integrity check fails and there is content missing. Use the first tax warning file as the second tax warning file, including: Obtain the number of pixel points in the text mutilation recognition area. If it is determined that the number of pixel points is empty, it is determined that the integrity check passes; If it is determined that the number of pixel points is not empty, obtain the coordinate values of the abscissa and ordinate of all pixel points. If the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, it is determined that the integrity check passes; If there are coordinate values of the abscissa and ordinate of pixel points that do not pass the regularity check, it is determined that there are mutilated characters in the text mutilation recognition area, and it is determined that the integrity check fails.
3. The automated tax monitoring and warning method according to claim 2, wherein The step of classifying the edge coordinate set to obtain four subsets of edge coordinate lines for the four edge lines, generating an extension line corresponding to each edge line based on the subset of edge coordinate lines, and obtaining the text mutilation recognition area based on the edge line and the corresponding extension line, includes: Classify all pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain four subsets of edge coordinate lines for the four edge lines.
4. The automated tax monitoring and warning method according to claim 3, wherein The step of classifying all pixel points in the edge coordinate set according to the abscissa value and the ordinate value to obtain four subsets of edge coordinate lines for the four edge lines, includes: Determine the interval of the abscissa of all pixel points, determine the maximum ordinate value of all pixel points in the interval of the abscissa to obtain the subset of edge coordinate lines for the first edge, and determine the minimum ordinate value of all pixel points in the interval of the abscissa to obtain the subset of edge coordinate lines for the second edge; Determine the interval of the ordinate of all pixel points, determine the maximum abscissa value of all pixel points in the interval of the ordinate to obtain the subset of edge coordinate lines for the third edge, and determine the minimum abscissa value of all pixel points in the interval of the ordinate to obtain the subset of edge coordinate lines for the fourth edge.
5. The automated tax monitoring and warning method according to claim 2, wherein The step of obtaining the coordinate values of the abscissa and ordinate of all pixel points, and if the coordinate values of the abscissa and ordinate of all pixel points pass the regularity check, determining that the integrity check passes, includes: Obtain all pixel points with a preset pixel value as the points to be verified; If it is determined that the ordinate values of the points to be verified at adjacent abscissas are the same, or the absolute value of the difference in ordinate values is the same preset value, then it is determined that the regular pattern verification is passed; If it is determined that the abscissa values of the points to be verified at adjacent ordinates are the same, or the absolute value of the difference in abscissa values is the same preset value, then it is determined that the regular pattern verification is passed; If it is determined that the point to be verified is less than the preset value, then it is determined that the regular pattern verification is passed.
6. The automated tax monitoring and warning method according to claim 1, wherein obtaining a target text comparison box by determining the standard text comparison boxes corresponding to the first difference and the second difference with the closest absolute value relationship, and performing filling analysis on the loss target after superposition based on the target text comparison box to obtain a corresponding filling result, including: determining the attribute of the text mutilation recognition area where the loss target is located to obtain the alignment method of the target text comparison box, and each attribute of the text mutilation recognition area has a preset alignment method.
7. The automated tax monitoring and warning method according to claim 6, characterized in that, It also includes: performing text comparison analysis by traversing the text database based on the position and glyph of the loss target within the target text comparison box, determining the text corresponding to the position and glyph within the target text comparison box as the filling text, and obtaining a filling result, including: performing a first extension process on the text mutilation recognition area based on the target text comparison box, and filling the filling text into the target text comparison box; If there are multiple filling texts, perform a second extension process on the corresponding target text comparison box to obtain a new target text comparison box and fill in all the filling texts in sequence, so that each loss target has a corresponding filling text; If the filling text of any loss target is empty, fill in an empty space within the corresponding target text comparison box.
8. The automated tax monitoring and warning method according to claim 1, wherein generating a monitoring and warning table after performing warning monitoring processing based on tax types and type values, and the monitoring and warning table at least includes information areas of different risk levels, including: initializing the monitoring and warning table, and the monitoring and warning table includes a low-risk area and a high-risk area; counting all tax types and type values that meet the requirements and filling them into the low-risk area, and correspondingly setting the corresponding contents in the first tax warning file and the second tax warning file with the information in the low-risk area; counting all tax types and type values that do not meet the requirements and filling them into the high-risk area, and correspondingly setting the corresponding contents in the first tax warning file and the second tax warning file with the information in the high-risk area.
9. The automated tax monitoring and warning method according to claim 8, wherein the step of counting all tax types and type values that do not meet the requirements and filling them into the high-risk area, and correspondingly setting the corresponding contents in the first tax warning file and the second tax warning file with the information in the high-risk area, includes: counting all tax types and type values that do not meet the requirements and filling them into the high-risk area; If it is determined that the characters of the tax types and type values that do not meet the requirements exist within the target text comparison box, add a filling verification label to the corresponding tax types and type values.
10. An automated tax monitoring and early warning system for the method according to any one of claims 1-9, characterized in that, It includes: The verification module is used to sequentially perform integrity verification on all the first tax warning files in the tax warning materials uploaded by the user. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file; The recognition module is used to perform edge contour recognition on the second tax warning file, determine the loss target superimposed on the edge contour, and perform filling analysis on the loss target based on the text database to obtain the corresponding filling result; The extraction module is used to extract the text information in the first tax warning file and the second tax warning file respectively, and determine the tax type and type value corresponding to each text information; The warning module is used to generate a monitoring warning table after performing warning monitoring processing based on the tax type and type value. The monitoring warning table at least includes information areas with different risk levels.
Citation Information
Patent Citations
Intelligent auditing method and device for inspection evidence material based on OCR (Optical Character Recognition) technology
CN114639173A
Electricity bill automatic data processing method and device and storage medium
CN116469120A