Automatic tax monitoring and early warning method and system
By performing integrity checks and edge profile identification on tax early warning materials, combined with the complementary analysis of the text database, a monitoring early warning table is generated, which solves the problems of inefficient processing of tax early warning materials and insufficient judgment of document integrity in the existing technology, and efficient and accurate tax early warning processing is achieved.
Patent Information
- Application Number
- CN202510484333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art is inefficient in the processing of tax warning materials and is prone to omissions, especially in judging file integrity and identifying incomplete text in images.
By verifying the integrity of the tax warning materials uploaded by users, identifying the image edge outline, determining the text defect area, and performing complementary analysis based on the text database, a monitoring warning table is generated to display information of different risk levels.
It realizes efficient and accurate processing of tax warning materials, improves the accuracy of document integrity verification, and ensures the accuracy of tax warning and data quality.
Smart Images

Figure CN120013695A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to automatic processing technology, and in particular to an automatic tax monitoring and early warning method and system. Background Art
[0002] In the prior art, some tax warning materials are processed manually, which is not only inefficient but also prone to omissions. Even if there are some automated processing methods, there are still many shortcomings. For example, in terms of integrity verification, it is difficult for the prior art to accurately determine whether there is any missing content in the file, especially for files containing image materials, it is impossible to effectively identify the incompleteness of the text in the image. The traditional processing method is difficult to meet the needs of efficient and accurate processing.
[0003] Therefore, there is an urgent need for an automated tax monitoring and early warning method to improve processing efficiency and accuracy. Summary of the invention
[0004] The embodiment of the present invention provides an automated tax monitoring and early warning method and system, which can provide an automated tax monitoring and early warning method to improve processing efficiency and accuracy.
[0005] A first aspect of an embodiment of the present invention provides an automated tax monitoring and early warning method, comprising: All first tax warning files in the tax warning materials uploaded by the user are checked for integrity in turn. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file; Perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform complement analysis on the loss target based on the text database to obtain the corresponding complement result; Extracting text information from the first tax warning file and the second tax warning file respectively, and determining the tax category and category value corresponding to each piece of text information; After early warning monitoring processing is performed based on the tax types and type values, a monitoring and early warning table is generated, and the monitoring and early warning table at least includes information areas of different risk levels.
[0006] Optionally, in a possible implementation of the first aspect, the step of sequentially verifying the integrity of all first tax warning files in the tax warning materials uploaded by the user, and using the first tax warning file as the second tax warning file if it is determined that there is content missing, includes: If it is determined that the tax warning materials uploaded by the user contain image materials, the coordinates of all edge pixel points in the image materials are obtained to obtain an edge coordinate set; The edge coordinate set is classified to obtain four edge coordinate line subsets, and the extension line corresponding to each edge line is generated based on the edge coordinate line subsets, and the middle area formed by the edge line and the corresponding extension line is used as the text incomplete recognition area; If it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check has not passed and there is content missing, and the first tax warning file will be used as the second tax warning file.
[0007] Optionally, in a possible implementation manner of the first aspect, classifying the edge coordinate set to obtain edge coordinate line subsets of four edge lines, generating an extension line corresponding to each edge line based on the edge coordinate line subsets, and obtaining the incomplete text recognition area based on the edge line and the corresponding extension line, includes: All pixel points in the edge coordinate set are classified according to the horizontal coordinate value and the vertical coordinate value to obtain the edge coordinate line subsets of 4 edge lines; Determining an extension direction corresponding to each edge line, wherein each category of edge lines has a preset extension direction; According to each pixel point in the edge line, a preset point position is extended according to the extension direction to determine the corresponding extension point, and the adjacent extension points corresponding to each edge line are sequentially connected to generate the corresponding extension line.
[0008] Optionally, in a possible implementation manner of the first aspect, all pixel points in the edge coordinate set are classified according to abscissa values and ordinate values to obtain edge coordinate line subsets of four edge lines, including: Determine the interval of the horizontal coordinates of all the pixels, determine the maximum vertical coordinate value of all the pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the first edge, and determine the minimum vertical coordinate value of all the pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the second edge; Determine the interval of the ordinates of all pixels, determine the maximum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the third edge, and determine the minimum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the fourth edge.
[0009] Optionally, in a possible implementation of the first aspect, if it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check fails and there is content missing, and the first tax warning file is used as the second tax warning file, including: Get the number of pixels in the incomplete text recognition area. If the number of pixels is empty, the integrity check is passed. If it is determined that the number of pixel points is not empty, the coordinate values of the horizontal and vertical coordinates of all pixel points are obtained. If the coordinate values of the horizontal and vertical coordinates of all pixel points pass the regularity verification, it is determined that the integrity verification has passed; If the coordinate values of the horizontal and vertical coordinates of all the pixels fail to pass the regularity check, it is determined that incomplete characters exist in the incomplete character recognition area, and it is determined that the integrity check fails.
[0010] Optionally, in a possible implementation manner of the first aspect, obtaining coordinate values of abscissas and ordinates of all pixel points, and if the coordinate values of abscissas and ordinates of all pixel points pass the regularity verification, then judging that the integrity verification has passed includes: Obtain all pixel points with preset pixel values as points to be verified; If the ordinate values of the adjacent horizontal coordinates of the points to be verified are the same, or the absolute value of the difference between the ordinate values is the same preset value, then it is judged that the regularity verification has been passed; If the horizontal coordinate values of the adjacent vertical coordinate points to be verified are the same, or the absolute value of the horizontal coordinate value difference is the same preset value, then it is judged that the regularity verification has been passed; If the point to be verified is judged to be less than the preset value, it is judged to pass the regularity verification.
[0011] Optionally, in a possible implementation of the first aspect, the performing edge contour recognition on the second tax early warning file, determining a loss target superimposed with the edge contour, and performing a completion analysis on the loss target based on a text database to obtain a corresponding completion result includes: Obtain the non-edge outline text in the second tax warning file, and select the center point of a non-edge text in each line as the alignment point; Generate a preset standard text comparison frame for non-edge text. If there are multiple standard text comparison frames, align the center points of the standard text comparison frames with the alignment points in sequence, and determine the first edge horizontal coordinates of the text on both sides corresponding to the alignment point close to the alignment point, and the first edge horizontal coordinates are 2; Obtaining the second edge horizontal coordinates of both sides of the standard text comparison frame, the second edge horizontal coordinates being 2, and obtaining the third edge horizontal coordinates of both sides of the standard text, the third edge horizontal coordinates being 2; Based on the numerical relationship between the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate, the standard text comparison frame is screened to obtain the target text comparison frame, and based on the target text comparison frame, the lost target is superimposed and then a completion analysis is performed to obtain the corresponding completion result.
[0012] Optionally, in a possible implementation manner of the first aspect, the method of screening the standard text comparison frame based on the numerical relationship among the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result includes: Taking the center point as the dividing point, the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate are divided toward both sides to obtain two sets of edge horizontal coordinate sets; Calculate the difference between the first edge horizontal coordinate and the second edge horizontal coordinate in each set of edge horizontal coordinates to obtain a first difference, and calculate the difference between the second edge horizontal coordinate and the third edge horizontal coordinate to obtain a second difference; The standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship is determined to obtain the target text comparison frame, and the loss target is superimposed based on the target text comparison frame and then a completion analysis is performed to obtain the corresponding completion result.
[0013] Optionally, in a possible implementation manner of the first aspect, determining a standard text comparison frame corresponding to the first difference and the second difference having the closest absolute value relationship to obtain a target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result, including: Determine the attribute of the incomplete text recognition area where the loss target is located to obtain the alignment of the target text comparison frame, and the attribute of each incomplete text recognition area has a preset alignment; Based on the preset alignment method, the target text comparison frame and the loss target are superimposed to obtain the position and shape of the loss target in the target text comparison frame; Based on the position and glyph of the loss target in the target text comparison frame, the text database is traversed to perform text comparison analysis, the text corresponding to the position and glyph in the target text comparison frame is determined as the completion text, and the completion result is obtained.
[0014] Optionally, in a possible implementation of the first aspect, the method further includes: The step of traversing the text database to perform text comparison analysis based on the position and glyph of the lost target in the target text comparison frame, determining the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtaining the completion result includes: Based on the target text comparison frame, the incomplete text recognition area is extended once, and the completed text is filled into the target text comparison frame; If there are multiple complementing characters, a secondary extension process is performed on the corresponding target character comparison frame to obtain a new target character comparison frame and all the complementing characters are filled in sequentially, so that each loss target has a corresponding complementing character; If the completion text of any loss target is empty, fill in the blank in the corresponding target text comparison box.
[0015] Optionally, in a possible implementation manner of the first aspect, the monitoring and early warning table is generated after the early warning monitoring processing is performed based on the tax type and the type value, and the monitoring and early warning table at least includes information areas of different risk levels, including: Initialize a monitoring and warning table, wherein the monitoring and warning table includes low-risk areas and high-risk areas; Count all tax types and type values that meet the requirements and fill them into the low-risk area, and set the corresponding contents in the first tax warning file and the second tax warning file to correspond to the information in the low-risk area; All tax types and type values that do not meet the requirements are counted and filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set corresponding to the information in the high-risk area.
[0016] Optionally, in a possible implementation of the first aspect, the counting of all tax types and type values that do not meet the requirements and filling them into the high-risk area, and setting the corresponding contents in the first tax early warning file and the second tax early warning file to correspond to the information of the high-risk area, includes: Count all tax types and type values that do not meet the requirements and fill them into the high-risk area; If the characters of the tax type and type value that are judged not to be satisfied exist in the target text comparison box, a completion check tag is added to the corresponding tax type and type value.
[0017] A second aspect of an embodiment of the present invention provides an automated tax monitoring and early warning system, comprising: A verification module is used to sequentially verify the integrity of all first tax warning files in the tax warning materials uploaded by the user. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file; The recognition module is used to perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform a completion analysis on the loss target based on the text database to obtain a corresponding completion result; An extraction module, used to extract text information in the first tax warning file and the second tax warning file respectively, and determine the tax category and category value corresponding to each piece of text information; The early warning module is used to generate a monitoring early warning table after performing early warning monitoring based on the tax types and type values. The monitoring early warning table at least includes information areas of different risk levels.
[0018] Beneficial effects: The present invention can accurately determine whether there is any missing content in the file by performing an integrity check on the first tax warning file in the tax warning materials uploaded by the user. For files containing image materials, by obtaining the coordinates of all edge pixel points in the image material, classifying the edge coordinate set, determining the incomplete text recognition area, and further determining whether there is incomplete text in the area, an accurate check of the file integrity is achieved. By analyzing and processing the coordinates, the incomplete text area is located to determine whether the text is complete. If incomplete text is found in the incomplete text recognition area, the file is marked as the second tax warning file for subsequent targeted processing, which effectively avoids the problem of inaccurate tax warnings caused by missing file content and improves the quality of tax warning data.
[0019] When the first tax warning file is determined to have missing content and marked as the second tax warning file, the present invention can effectively restore the integrity of the file by performing edge contour recognition, determining the loss target superimposed with the edge contour, and completing the loss target based on the text database. Specifically, by obtaining the text of the non-edge contour in the second tax warning file, determining the alignment point, generating a standard text comparison frame, screening the target text comparison frame, and finally completing the text of the loss target, the text information in the file is complete and accurate, providing a reliable data basis for subsequent tax analysis.
[0020] The present invention extracts the text information in the first tax warning file and the second tax warning file respectively, determines the tax type and type value corresponding to each text information, and performs early warning monitoring based on this information to generate a monitoring early warning table containing information areas of different risk levels. The system can accurately count the tax types and values that meet the requirements and those that do not meet the requirements, fill them into low-risk areas and high-risk areas respectively, and set the corresponding content in the file to correspond to the regional information. At the same time, for tax information that does not meet the requirements and whose characters exist in the target text comparison box, a completion verification label is added to further improve the accuracy of early warning monitoring. The monitoring early warning table can clearly display the risk levels of different tax types and values. Tax management personnel can intuitively understand the tax risk situation, promptly discover high-risk tax information, and trace and analyze its source and specific situation, providing strong support for tax decision-making and effectively improving the efficiency and accuracy of tax management. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of an automated tax monitoring and early warning method provided by an embodiment of the present invention. Figure 2 A schematic diagram of an extension line provided by an embodiment of the present invention; Figure 3This is a structural diagram of an automated tax monitoring and early warning system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] See also Figure 1 , is a flow chart of an automated tax monitoring and early warning method provided by an embodiment of the present invention, the method comprising: S1, all first tax warning files in the tax warning materials uploaded by the user are sequentially verified for integrity. If it is determined that there is content missing, the first tax warning file is used as the second tax warning file.
[0024] In the automated tax monitoring and early warning system, the integrity check of the first tax warning file in the tax warning materials uploaded by the user is required to ensure the quality of the tax warning data and the accuracy of subsequent analysis and early warning. When it is found that the file has missing content, the file will be marked as the second tax warning file for subsequent targeted processing.
[0025] In some embodiments, the step of sequentially verifying the integrity of all first tax warning files in the tax warning materials uploaded by the user, and using the first tax warning file as the second tax warning file if it is determined that there is content missing, includes: S11, if it is determined that the tax warning materials uploaded by the user contain image materials, the coordinates of all edge pixel points in the image materials are obtained to obtain an edge coordinate set.
[0026] The image material can be an example of tax materials such as invoice materials. The edge refers to the edge of the image. When the system determines that the tax warning materials uploaded by the user contain image materials, it will obtain the coordinates of the edge pixels. Scan the image material and accurately extract the coordinates of all edge pixels in the image to form an edge coordinate set. The edge coordinate set is required for subsequent analysis of the image structure and recognition of the incomplete text area.
[0027] S12, classifying the edge coordinate set to obtain edge coordinate line subsets of four edge lines, generating an extension line corresponding to each edge line based on the edge coordinate line subsets, and using the middle area formed by the edge line and the corresponding extension line as the text incomplete recognition area.
[0028] See also Figure 2 The system will classify the acquired edge coordinate set to form 4 edge coordinate line subsets. Based on the subsets, the extension line corresponding to each edge line is generated. The edge line and its corresponding extension line are used to determine the text defect recognition area, providing a prerequisite for the subsequent accurate positioning of the text defect position.
[0029] The step of classifying the edge coordinate set to obtain edge coordinate line subsets of four edge lines, generating an extension line corresponding to each edge line based on the edge coordinate line subsets, and obtaining a text incomplete recognition area based on the edge line and the corresponding extension line includes: S121, classify all pixel points in the edge coordinate set according to the horizontal coordinate value and the vertical coordinate value to obtain edge coordinate line subsets of 4 edge lines.
[0030] To form the edge coordinate line subsets of the four edge lines, the system classifies all the pixels in the edge coordinate set according to the horizontal and vertical coordinate values. By combing the coordinate data, the pixels are classified into the corresponding edge line coordinate subsets to clearly define the different parts of the image edge.
[0031] The edge coordinate line subsets of four edge lines are obtained by classifying all pixel points in the edge coordinate set according to the horizontal coordinate value and the vertical coordinate value, including: S1211, determine the interval of the horizontal coordinates of all pixels, determine the maximum vertical coordinate value of all pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the first edge, and determine the minimum vertical coordinate value of all pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the second edge.
[0032] The system first determines the interval of the horizontal coordinates of all pixels, and within this interval, finds the maximum value of the vertical coordinate of the pixel points in the horizontal coordinate interval, and classifies these pixels into the edge coordinate line subset of the first edge, that is, the top line; finds the minimum value of the vertical coordinate of the pixel points in the horizontal coordinate interval, and classifies them into the edge coordinate line subset of the second edge, that is, the bottom line, thereby realizing the preliminary division of the image edge from the vertical dimension.
[0033] S1212, determine the interval of the ordinates of all pixels, determine the maximum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the third edge, and determine the minimum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the fourth edge.
[0034] The system determines the interval of the ordinates of all pixels, and within this interval, finds the maximum value of the horizontal coordinate of the pixel in the ordinate interval, and classifies it as the edge coordinate line subset of the third edge, that is, the rightmost line; finds the minimum value of the horizontal coordinate of the pixel in the ordinate interval, and classifies it as the edge coordinate line subset of the fourth edge, that is, the leftmost line, further refining the division of the image edge from the horizontal dimension.
[0035] S122, determining an extension direction corresponding to each edge line, wherein each category of edge lines has a preset extension direction.
[0036] The extension direction is preset, for example, the top edge line moves downward and extends, the bottom line moves upward and extends, the leftmost edge line moves right and extends, and the rightmost line moves left and extends.
[0037] S123, extending a preset point position according to an extension direction for each pixel point in the edge line, determining a corresponding extension point, and sequentially connecting adjacent extension points corresponding to each edge line to generate a corresponding extension line.
[0038] The system moves each pixel point in the edge line along the extension direction to determine the corresponding extension point. The adjacent extension points corresponding to each edge line are connected in sequence to generate the corresponding extension line, thereby constructing the text incomplete recognition area to provide a guarantee for the subsequent accurate recognition of text incompleteness.
[0039] S13, if it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check has not passed and there are missing contents, and the first tax warning file is used as the second tax warning file.
[0040] The system determines whether there are incomplete texts in the area to further determine whether the file has passed the integrity check. If it is determined that there are incomplete texts, it means that the file has not passed the integrity check and there is content missing, and the first tax warning file is marked as the second tax warning file for subsequent special processing; if there are no incomplete texts, it is considered that the file has passed the integrity check and can enter the subsequent normal tax warning analysis process.
[0041] If it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check has not been passed and there is content missing, and the first tax warning file is used as the second tax warning file, including: S131, obtaining the number of pixel points in the incomplete text recognition area, and if it is determined that the number of pixel points is empty, it is determined that the integrity verification has passed.
[0042] The system first obtains the number of pixels in the incomplete text recognition area. If the number of pixels in the area is empty, that is, there are no pixels, this means that no text-related information is detected in the area, and it can be preliminarily determined that there is no incomplete text in the area, thereby determining that the first tax warning file has passed the integrity check. Because under normal circumstances, if the text is complete, a certain number of pixels will not be detected in the recognition area to form the text. If the number of pixels is empty, it is highly likely that the text is complete and there is no incompleteness.
[0043] S132, if it is determined that the number of pixel points is not empty, the coordinate values of the horizontal and vertical coordinates of all pixel points are obtained. If the coordinate values of the horizontal and vertical coordinates of all pixel points pass the regularity verification, it is determined that the integrity verification is passed.
[0044] When it is determined that the number of pixels in the text defect recognition area is not empty, the system further obtains the coordinate values of the horizontal and vertical coordinates of all pixels and performs a regularity check on these coordinate values to more accurately determine whether the text is defective. If the coordinate values of the horizontal and vertical coordinates of all pixels pass the regularity check, it is determined that the first tax warning file passes the integrity check.
[0045] The step of obtaining the coordinate values of the horizontal and vertical coordinates of all the pixels, and judging that the integrity check has been passed if the coordinate values of the horizontal and vertical coordinates of all the pixels pass the regularity check, includes: S1321, obtaining all pixel points with preset pixel values as points to be verified.
[0046] The system first obtains all pixels with preset pixel values as points to be verified. The preset pixel values are set based on the display characteristics of text in the image and past experience. These pixels with specific pixel values are more likely to be key pixels that constitute the text. For example, the pixel value corresponding to black.
[0047] S1322: If it is determined that the ordinate values of the adjacent horizontal coordinate points to be verified are the same, or the absolute value of the difference between the ordinate values is the same preset value, it is determined that the regularity verification has been passed.
[0048] This step is described for the incomplete text recognition areas above and below.
[0049] The system determines the ordinate values of adjacent horizontal coordinate points to be verified. If the ordinate values of adjacent horizontal coordinate points to be verified are the same, such as a horizontal line, or the absolute value of the difference between the ordinate values is the same preset value, such as a broken line.
[0050] The above shows that these points to be verified show a certain regularity in the vertical direction. It can be preliminarily judged that they are not in the form of text, but in the form of horizontal lines or broken lines that are often found in some documents. For example, invoice documents often have wireframes.
[0051] S1323, if it is determined that the horizontal coordinate values of the adjacent vertical coordinate points to be verified are the same, or the absolute value of the difference between the horizontal coordinate values is the same preset value, it is determined that the regularity verification has been passed; This step is described corresponding to the left and right incomplete text recognition areas.
[0052] The system determines the horizontal coordinate values of adjacent vertical coordinate points to be verified. If the horizontal coordinate values of adjacent vertical coordinate points to be verified are the same, for example, in the form of a vertical line, or the absolute value of the horizontal coordinate value difference is the same preset value, for example, in the form of a broken line.
[0053] The above shows that these points to be verified show a certain regularity in the vertical direction. It can be preliminarily judged that they are not in the form of text, but in the form of horizontal lines or broken lines that are often found in some documents. For example, invoice documents often have wireframes.
[0054] S1324: If the point to be verified is determined to be less than the preset value, it is determined that the regularity verification has been passed.
[0055] The system determines whether the number of points to be verified is less than the preset value. If it is determined that the number of points to be verified is less than the preset value, it means that the number of key pixel points constituting the text in the area is small, and the regularity verification is passed.
[0056] S133, if there is a pixel point whose horizontal and vertical coordinate values fail to pass the regularity check, it is determined that there is incomplete text in the incomplete text recognition area, and it is determined that it fails to pass the integrity check.
[0057] If, during the regularity check of the horizontal and vertical coordinate values of the pixel points, there are coordinate values that fail to pass the regularity check, the system determines that there are incomplete characters in the incomplete character recognition area, and further determines that the first tax warning file fails the integrity check. This means that the characters in the image material in the file are incomplete, which may affect the accuracy of the tax warning analysis, so the file is marked as the second tax warning file.
[0058] S2, performing edge contour recognition on the second tax warning file, determining the loss target superimposed with the edge contour, and performing completion analysis on the loss target based on the text database to obtain a corresponding completion result.
[0059] When the first tax warning file is judged to have missing content and marked as the second tax warning file, the system will perform further processing on it. First, the second tax warning file is subjected to edge contour recognition to determine the edge contour of the image part. On this basis, the loss targets superimposed on these edge contours are determined. These loss targets are usually the missing information parts due to incomplete text. Subsequently, the system performs a completion analysis on the loss targets based on a pre-established text database. The text database stores the features and information of various standard texts. The system matches and analyzes the loss targets with the text in the database to obtain the corresponding completion results, thereby restoring the missing text information in the file and improving the integrity and accuracy of the tax warning file.
[0060] In some embodiments, the performing of edge contour recognition on the second tax early warning file, determining the loss target superimposed with the edge contour, and performing a completion analysis on the loss target based on a text database to obtain a corresponding completion result includes: S21, obtaining the non-edge outline text in the second tax warning file, and selecting the center point of a non-edge text in each line as the alignment point.
[0061] The system extracts the text content without edge contours from the second tax warning file. For each line of non-edge text, the system selects the center point of the text as the alignment point. By determining the alignment point, a reference point is provided for the subsequent generation of a standard text comparison frame, so that subsequent text analysis and processing can be carried out more accurately, which helps to improve the accuracy of the completion analysis.
[0062] Among them, the method for obtaining the center point of the text can be to determine the coordinates of its leftmost point and rightmost point, determine the horizontal coordinate value through the coordinates of the leftmost point and the rightmost point, then determine the coordinates of its topmost point and the bottommost point, determine the vertical coordinate value through the coordinates of the topmost point and the bottommost point, and combine the horizontal coordinate value and the vertical coordinate value to obtain the center point.
[0063] S22, generate a preset standard text comparison frame for non-edge text. If there are multiple standard text comparison frames, align the center points of the standard text comparison frames with the alignment points in turn, and determine the first edge horizontal coordinates of the text on both sides corresponding to the alignment point close to the alignment point. The first edge horizontal coordinates are 2.
[0064] The system generates a preset standard text comparison box for non-edge text. If multiple standard text comparison boxes are generated, the system will align the center point of each standard text comparison box with the previously determined alignment point in turn. After alignment, the system determines the first edge horizontal coordinates of the text on both sides of the alignment point close to the alignment point, one first edge horizontal coordinate in each direction, for a total of two. These first edge horizontal coordinates reflect the horizontal boundary position of the text related to the alignment point.
[0065] S23, obtaining the second edge horizontal coordinates of both sides of the standard text comparison frame, the second edge horizontal coordinates are 2, and obtaining the third edge horizontal coordinates of both sides of the standard text, the third edge horizontal coordinates are 2.
[0066] The system further obtains the second edge horizontal coordinates on both sides of the standard text comparison box, one in each direction, for a total of two. These two second edge horizontal coordinates determine the horizontal boundary range of the standard text comparison box itself. At the same time, the system obtains the third edge horizontal coordinates on both sides of the standard text, one in each direction, for a total of two. The third edge horizontal coordinate provides more detailed position information about the standard text in the horizontal direction. By obtaining these different levels of edge horizontal coordinates, the system can fully understand the spatial position and range of the standard text comparison box and the standard text in the horizontal direction.
[0067] S24, based on the numerical relationship between the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate, the standard text comparison frame is screened to obtain a target text comparison frame, and based on the target text comparison frame, the lost target is superimposed and a completion analysis is performed to obtain a corresponding completion result.
[0068] After obtaining the first edge horizontal coordinate, the second edge horizontal coordinate and the third edge horizontal coordinate, the system screens the standard text comparison frame based on the numerical relationship between these horizontal coordinates to determine the most suitable target text comparison frame. After determining the target text comparison frame, it is superimposed with the loss target, and a detailed completion analysis of the loss target is performed based on the text database to obtain the corresponding completion result and restore the missing text information in the second tax warning file.
[0069] The method of screening the standard text comparison frame based on the numerical relationship among the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain the corresponding completion result includes: S241, taking the center point as a dividing point, dividing the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate toward both sides to obtain two sets of edge horizontal coordinate sets.
[0070] The system uses the previously determined center point (such as the alignment point) as the dividing point to divide the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate to the two sides. In this way, these horizontal coordinates are divided into two sets of edge horizontal coordinates. This grouping operation helps to analyze and compare the horizontal coordinates on both sides separately in the future, making the screening of the standard text comparison box more detailed and accurate.
[0071] S242, calculating the difference between the first edge horizontal coordinate and the second edge horizontal coordinate in each set of edge horizontal coordinates to obtain a first difference, and calculating the difference between the second edge horizontal coordinate and the third edge horizontal coordinate to obtain a second difference.
[0072] For each set of edge horizontal coordinates, the system calculates the difference between the first edge horizontal coordinate and the second edge horizontal coordinate in the set to obtain the first difference. At the same time, the difference between the second edge horizontal coordinate and the third edge horizontal coordinate is calculated to obtain the second difference. These differences reflect the distance relationship between different edge horizontal coordinates. By calculating these differences, the characteristics of the standard text comparison box in the horizontal direction can be quantified. The calculation results of these differences will serve as an important basis for the subsequent screening of target text comparison boxes, helping the system to determine which standard text comparison box best matches the actual text layout in the file.
[0073] S243, determining the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result.
[0074] The system compares the absolute value relationship between the first difference and the second difference corresponding to all standard text comparison frames, finds the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship, and determines it as the target text comparison frame. This target text comparison frame is considered to be the most consistent with the actual text layout and loss target in the file. After determining the target text comparison frame, the system superimposes the loss target based on the target text comparison frame and performs a completion analysis to obtain the corresponding completion result.
[0075] The step of determining the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain the corresponding completion result includes: S2431, determining the attributes of the incomplete text recognition region where the loss target is located to obtain the alignment of the target text comparison frame, and the attributes of each incomplete text recognition region have a preset alignment.
[0076] The system first determines the properties of the text defect recognition area where the loss target is located. The properties of each text defect recognition area are pre-set with the corresponding alignment. By determining the properties of the area where the loss target is located, the system can obtain the alignment of the target text comparison box. This alignment determination helps to accurately superimpose the target text comparison box with the loss target in the future, ensuring the accuracy of the completion analysis.
[0077] The attribute of each incomplete text recognition area having a preset alignment means that when the incomplete text is located in different areas, the incomplete position may be different. For example, for the incomplete text in the left incomplete text recognition area, the missing part is generally the left part of the text; for the incomplete text in the right incomplete text recognition area, the missing part is generally the right part of the text, and so on.
[0078] S2432: superimpose the target text comparison frame and the loss target based on the preset alignment method to obtain the position and font shape of the loss target in the target text comparison frame.
[0079] Based on the preset alignment method, the system overlays the target text comparison frame with the loss target. The system obtains the specific position information of the loss target in the target text comparison frame, as well as the glyph features presented by the loss target. These position and glyph information are key data for subsequent text comparison analysis, which can help the system accurately find text that matches the loss target in the text database.
[0080] S2433, based on the position and glyph of the lost target in the target text comparison frame, traverse the text database to perform text comparison analysis, determine the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtain the completion result.
[0081] After obtaining the position and glyph information of the loss target in the target text comparison box, the system traverses the text database based on this information and performs detailed text comparison analysis. The purpose is to find the text that matches the position and glyph in the target text comparison box, and determine it as the completion text. Finally, accurate completion results are obtained to complete the text completion work of the loss target in the second tax warning file.
[0082] The step of traversing the text database to perform text comparison analysis based on the position and glyph of the lost target in the target text comparison frame, determining the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtaining the completion result includes: S24331, performing an extension process on the incomplete text recognition area based on the target text comparison frame, and filling the completed text into the target text comparison frame; The system first performs an extension process on the incomplete text recognition area based on the target text comparison frame. This extension process is to expand the scope of the target text comparison frame so that it can better accommodate possible complementary text. After completing the extension process, the system fills the determined complementary text into the target text comparison frame. In this way, the text complement of the lost target is initially achieved, making the text information in the target text comparison frame more complete.
[0083] S24332, if there are multiple complementing characters, a secondary extension process is performed on the corresponding target character comparison frame to obtain a new target character comparison frame and all the complementing characters are filled in sequentially, so that each loss target has a corresponding complementing character; When there are multiple complementary characters, it means that the loss target may have multiple matching characters. At this time, the system performs secondary extension processing on the corresponding target character comparison box to generate a new target character comparison box. The secondary extension processing is to further expand the scope of the target character comparison box to meet the filling requirements of multiple complementary characters.
[0084] S24333, if the completion text of any loss target is empty, fill in the blank in the corresponding target text comparison box.
[0085] If after traversing the text database for text comparison analysis, it is found that the complementary text of any lost target is empty, that is, no text corresponding to the position and glyph of the lost target is found, the system will fill in the blank in the corresponding target text comparison box. This processing method ensures the consistency and accuracy of the data, clearly identifies the situation where the lost target cannot be completed through the existing text database, and provides clear information for subsequent processing. For example, if there is no matching text in the text database, the system will fill in the blank in the corresponding target text comparison box so that the staff can discover it in time and take further measures, such as manual review or supplementing the text database.
[0086] S3, extracting text information from the first tax warning file and the second tax warning file respectively, and determining the tax type and type value corresponding to each piece of text information.
[0087] The system processes the first tax warning file and the second tax warning file respectively, and extracts the text information in the file through text recognition technology. After extracting the text information, the system further analyzes each text message and determines the corresponding tax type and type value according to the pre-set rules and standards. For example, for text information involving amounts, the system will determine which tax item it belongs to (such as value-added tax, income tax, etc.) and the specific amount value. Through this step, the system converts the text information in the tax file into structured data that can be used for early warning monitoring, providing basic data support for the subsequent generation of monitoring and early warning tables.
[0088] S4, generating a monitoring and warning table after performing early warning monitoring based on the tax types and type values, the monitoring and warning table at least including information areas of different risk levels.
[0089] After obtaining the tax types and tax values, the system performs early warning monitoring based on this information and finally generates a monitoring early warning table. The monitoring early warning table is an important output of tax early warning analysis, which at least includes information areas of different risk levels so that users can intuitively understand the tax risk situation.
[0090] In some embodiments, the monitoring and early warning table is generated after the early warning monitoring process based on the tax type and the type value, and the monitoring and early warning table includes at least information areas of different risk levels, including: S41, initializing a monitoring and warning table, wherein the monitoring and warning table includes low-risk areas and high-risk areas.
[0091] The system first initializes the monitoring and early warning table. During the initialization process, the monitoring and early warning table is set to include at least two basic information areas: low-risk area and high-risk area. The low-risk area is used to display tax information that meets the requirements and has relatively low risks; the high-risk area is used to display tax information that does not meet the requirements and may have greater tax risks. This regional division method provides a framework for the subsequent classification and display of tax information.
[0092] S42, counting all tax types and type values that meet the requirements and filling them into the low-risk area, and setting the corresponding contents in the first tax warning file and the second tax warning file to correspond to the information of the low-risk area.
[0093] The system collects statistics on all tax types and type values, and selects tax information that meets the requirements. "Meeting the requirements" here is judged based on pre-set standards, such as the compliance of tax types, whether the type values are within the normal range, etc. The selected tax types and type values that meet the requirements are filled into the low-risk area of the monitoring warning table. At the same time, the system sets the corresponding content in the first tax warning file and the second tax warning file to correspond to the information in the low-risk area, so that users can clearly see the source and location of these low-risk tax information in the original file. In this way, users can quickly understand which tax information is in a low-risk state and the specific situation of the relevant information in the file.
[0094] S43, counting all tax types and type values that do not meet the requirements and filling them into the high-risk area, and setting the corresponding contents in the first tax warning file and the second tax warning file to correspond to the information of the high-risk area.
[0095] The system once again counts the tax types and type values, and this time filters out the tax information that does not meet the requirements. These situations that do not meet the requirements may include incorrect tax types, abnormal type values, etc. The filtered tax types and type values that do not meet the requirements are filled into the high-risk area of the monitoring and early warning table. Similarly, the system sets the corresponding content in the first tax early warning file and the second tax early warning file to correspond to the information in the high-risk area, so that users can accurately trace the source and specific circumstances of high-risk tax information. In this way, users can quickly pay attention to the parts that may have tax risks, and further analyze and process these high-risk information. Finally, the monitoring and early warning table generated through the above steps provides tax management personnel with an intuitive and clear tax risk monitoring tool, which helps to discover and solve tax problems in a timely manner.
[0096] The statistics of all tax types and tax type values that do not meet the requirements are filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set correspondingly to the information of the high-risk area, including: S431, count all tax types and type values that do not meet the requirements and fill them into the high-risk area.
[0097] The system conducts a comprehensive statistics of all tax types and type values, and determines which tax types and values do not meet the requirements based on pre-set standards. After screening out all tax types and type values that are judged to not meet the requirements, the system fills them into the high-risk area of the monitoring and early warning table. In this way, the high-risk area centrally displays all information that may have tax risks, allowing tax management personnel to quickly locate and focus on these key data.
[0098] S432: If it is determined that characters of the tax type and type value that are not satisfied exist in the target text comparison box, a completion check tag is added to the corresponding tax type and type value.
[0099] After filling the tax types and type values that do not meet the requirements into the high-risk area, the system further checks this information. For each tax type and type value that does not meet the requirements, the system determines whether its characters exist in the target text comparison box generated in the previous processing process. The target text comparison box is generated when the second tax warning file is completed and analyzed, and is used to determine and process the incomplete or missing parts of the text. If the characters of the tax types and type values that do not meet the requirements are determined to exist in the target text comparison box, it means that the tax information may not meet the requirements due to incomplete or missing text. In order to further verify and handle this situation, the system will add a completion verification tag to the corresponding tax types and type values. After adding the completion verification tag, tax management personnel can more clearly identify which high-risk tax information may be related to text processing, so as to conduct a targeted re-check of this information to see if there is inaccurate or incomplete text completion, so as to ensure the accuracy and completeness of tax information and further reduce tax risks.
[0100] See also Figure 3 , is a schematic diagram of the structure of an automated tax monitoring and early warning system provided by an embodiment of the present invention, the system comprising: A verification module is used to sequentially verify the integrity of all first tax warning files in the tax warning materials uploaded by the user. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file; The recognition module is used to perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform a completion analysis on the loss target based on the text database to obtain a corresponding completion result; An extraction module, used to extract text information in the first tax warning file and the second tax warning file respectively, and determine the tax category and category value corresponding to each piece of text information; The early warning module is used to generate a monitoring early warning table after performing early warning monitoring based on the tax types and type values. The monitoring early warning table at least includes information areas of different risk levels.
[0101] The present invention also provides a storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the methods provided by the various embodiments described above.
[0102] Among them, the storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also exist in a communication device as discrete components. The storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0103] The present invention also provides a program product, which includes an execution instruction, which is stored in a storage medium. At least one processor of a device can read the execution instruction from the storage medium, and at least one processor executes the execution instruction so that the device implements the methods provided in the above various embodiments.
[0104] In the above-mentioned terminal or server embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The automated tax monitoring and early warning method is characterized by: include: All first tax warning files in the tax warning materials uploaded by the user are checked for integrity in turn. If it is determined that there is any missing content, the first tax warning file will be used as the second tax warning file; Perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform complement analysis on the loss target based on the text database to obtain the corresponding complement result; Extract the text information in the first tax warning file and the second tax warning file respectively, and determine the tax type and type value corresponding to each text information; After early warning monitoring and processing based on tax types and type values, a monitoring and early warning table is generated, and the monitoring and early warning table at least includes information areas of different risk levels.
2. The automated tax monitoring and early warning method according to claim 1 is characterized in that: The integrity verification of all first tax warning files in the tax warning materials uploaded by the user is performed in sequence, and if it is determined that there is content missing, the first tax warning file is used as the second tax warning file, including: If it is determined that the tax warning materials uploaded by the user contain image materials, the coordinates of all edge pixel points in the image materials are obtained to obtain an edge coordinate set; The edge coordinate set is classified to obtain four edge coordinate line subsets, and the extension line corresponding to each edge line is generated based on the edge coordinate line subsets, and the middle area formed by the edge line and the corresponding extension line is used as the text incomplete recognition area; If it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check has not passed and there is content missing, and the first tax warning file will be used as the second tax warning file.
3. The automated tax monitoring and early warning method according to claim 2 is characterized in that: The edge coordinate set is classified to obtain edge coordinate line subsets of four edge lines, an extension line corresponding to each edge line is generated based on the edge coordinate line subsets, and a text incomplete recognition area is obtained based on the edge line and the corresponding extension line, including: All pixel points in the edge coordinate set are classified according to the horizontal coordinate value and the vertical coordinate value to obtain the edge coordinate line subsets of 4 edge lines; Determining an extension direction corresponding to each edge line, wherein each category of edge lines has a preset extension direction; According to each pixel point in the edge line, a preset point position is extended according to the extension direction to determine the corresponding extension point, and the adjacent extension points corresponding to each edge line are sequentially connected to generate the corresponding extension line.
4. The automated tax monitoring and early warning method according to claim 3 is characterized in that: The edge coordinate line subsets of four edge lines are obtained by classifying all pixel points in the edge coordinate set according to the horizontal coordinate value and the vertical coordinate value, including: Determine the interval of the horizontal coordinates of all the pixels, determine the maximum vertical coordinate value of all the pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the first edge, and determine the minimum vertical coordinate value of all the pixels in the interval of the horizontal coordinates to obtain the edge coordinate line subset of the second edge; Determine the interval of the ordinates of all pixels, determine the maximum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the third edge, and determine the minimum abscissa value of all pixels in the interval of the ordinates to obtain the edge coordinate line subset of the fourth edge.
5. The automated tax monitoring and early warning method according to claim 2 is characterized in that: If it is determined that there are incomplete characters in the incomplete character recognition area, it is determined that the integrity check has not been passed and there is content missing, and the first tax warning file is used as the second tax warning file, including: Get the number of pixels in the incomplete text recognition area. If the number of pixels is empty, the integrity check is passed. If it is determined that the number of pixel points is not empty, the coordinate values of the horizontal and vertical coordinates of all pixel points are obtained. If the coordinate values of the horizontal and vertical coordinates of all pixel points pass the regularity verification, it is determined that the integrity verification has passed; If the coordinate values of the horizontal and vertical coordinates of a pixel point fail to pass the regularity check, it is determined that incomplete text exists in the incomplete text recognition area and that the integrity check fails.
6. The automated tax monitoring and early warning method according to claim 5 is characterized in that: The method of obtaining the coordinate values of the horizontal and vertical coordinates of all the pixels, and judging that the coordinate values of the horizontal and vertical coordinates of all the pixels pass the regularity verification, includes: Obtain all pixel points with preset pixel values as points to be verified; If the ordinate values of the adjacent horizontal coordinates of the points to be verified are the same, or the absolute value of the difference between the ordinate values is the same preset value, then it is judged that the regularity verification has been passed; If the horizontal coordinate values of the adjacent vertical coordinate points to be verified are the same, or the absolute value of the horizontal coordinate value difference is the same preset value, then it is judged that the regularity verification has been passed; If the point to be verified is judged to be less than the preset value, it is judged to pass the regularity verification.
7. The automated tax monitoring and early warning method according to claim 5 is characterized in that: The performing edge contour recognition on the second tax early warning file, determining the loss target superimposed with the edge contour, and performing a completion analysis on the loss target based on the text database to obtain a corresponding completion result includes: Obtain the non-edge outline text in the second tax warning file, and select the center point of a non-edge text in each line as the alignment point; Generate a preset standard text comparison frame for non-edge text. If there are multiple standard text comparison frames, align the center points of the standard text comparison frames with the alignment points in sequence, and determine the first edge horizontal coordinates of the text on both sides corresponding to the alignment point close to the alignment point, and the first edge horizontal coordinates are 2; Obtaining the second edge horizontal coordinates of both sides of the standard text comparison frame, the second edge horizontal coordinates being 2, and obtaining the third edge horizontal coordinates of both sides of the standard text, the third edge horizontal coordinates being 2; Based on the numerical relationship between the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate, the standard text comparison frame is screened to obtain the target text comparison frame, and based on the target text comparison frame, the lost target is superimposed and then a completion analysis is performed to obtain the corresponding completion result.
8. The automated tax monitoring and early warning method according to claim 7 is characterized in that: The method of screening the standard text comparison frame based on the numerical relationship among the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result includes: Taking the center point as the dividing point, the first edge horizontal coordinate, the second edge horizontal coordinate, and the third edge horizontal coordinate are divided toward both sides to obtain two sets of edge horizontal coordinate sets; Calculate the difference between the first edge horizontal coordinate and the second edge horizontal coordinate in each set of edge horizontal coordinates to obtain a first difference, and calculate the difference between the second edge horizontal coordinate and the third edge horizontal coordinate to obtain a second difference; The standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship is determined to obtain the target text comparison frame, and the loss target is superimposed based on the target text comparison frame and then a completion analysis is performed to obtain the corresponding completion result.
9. The automated tax monitoring and early warning method according to claim 8 is characterized in that: The step of determining the standard text comparison frame corresponding to the first difference and the second difference with the closest absolute value relationship to obtain the target text comparison frame, and performing a completion analysis on the lost target after superimposing the target text comparison frame to obtain a corresponding completion result includes: Determine the properties of the incomplete text recognition area where the loss target is located to obtain the alignment of the target text comparison frame, and the properties of each incomplete text recognition area have a preset alignment; Based on the preset alignment method, the target text comparison frame and the loss target are superimposed to obtain the position and shape of the loss target in the target text comparison frame; Based on the position and glyph of the loss target in the target text comparison frame, the text database is traversed to perform text comparison analysis, the text corresponding to the position and glyph in the target text comparison frame is determined as the completion text, and the completion result is obtained.
10. The automated tax monitoring and early warning method according to claim 9 is characterized in that: Also includes: The step of traversing the text database to perform text comparison analysis based on the position and glyph of the lost target in the target text comparison frame, determining the text corresponding to the position and glyph in the target text comparison frame as the completion text, and obtaining the completion result includes: Based on the target text comparison frame, the incomplete text recognition area is extended once, and the completed text is filled into the target text comparison frame; If there are multiple complementing characters, a secondary extension process is performed on the corresponding target character comparison frame to obtain a new target character comparison frame and all the complementing characters are filled in sequentially, so that each loss target has a corresponding complementing character; If the completion text of any loss target is empty, fill in the blank in the corresponding target text comparison box.
11. The automated tax monitoring and early warning method according to claim 1, characterized in that: After the early warning monitoring process based on the tax type and the type value is performed, a monitoring early warning table is generated, and the monitoring early warning table at least includes information areas of different risk levels, including: Initialize a monitoring and warning table, wherein the monitoring and warning table includes low-risk areas and high-risk areas; Count all tax types and type values that meet the requirements and fill them into the low-risk area, and set the corresponding contents in the first tax warning file and the second tax warning file to correspond to the information in the low-risk area; All tax types and type values that do not meet the requirements are counted and filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set corresponding to the information in the high-risk area.
12. The automated tax monitoring and early warning method according to claim 11, characterized in that: The statistics of all tax types and tax type values that do not meet the requirements are filled into the high-risk area, and the corresponding contents in the first tax warning file and the second tax warning file are set correspondingly to the information of the high-risk area, including: Count all tax types and type values that do not meet the requirements and fill them into the high-risk area; If the characters of the tax type and type value that are judged not to be satisfied exist in the target text comparison box, a completion check tag is added to the corresponding tax type and type value.
13. The automated tax monitoring and early warning system is characterized by: include: A verification module is used to sequentially verify the integrity of all first tax warning files in the tax warning materials uploaded by the user. If it is determined that there is content missing, the first tax warning file will be used as the second tax warning file; The recognition module is used to perform edge contour recognition on the second tax early warning file, determine the loss target superimposed on the edge contour, and perform a completion analysis on the loss target based on the text database to obtain a corresponding completion result; An extraction module, used to extract text information in the first tax warning file and the second tax warning file respectively, and determine the tax category and category value corresponding to each piece of text information; The early warning module is used to generate a monitoring early warning table after performing early warning monitoring based on the tax types and type values. The monitoring early warning table at least includes information areas of different risk levels.
Citation Information
Patent Citations
Intelligent auditing method and device for inspection evidence material based on OCR (Optical Character Recognition) technology
CN114639173A
Electricity bill automatic data processing method and device and storage medium
CN116469120A
Invoice item name and tax classification verification method and related device
CN119091456A
Document processing apparatus
JP2011070529A
Methods systems and articles of manufacture for providing tax document guidance during preparation of electronic tax return
US9916627B1