Value-added tax bill tilt correction method, system, device and medium

By introducing a deep separable convolution module and a global attention module in the YOLOv7 network, the table linear detection and tilt angle correction of VAT bills is solved, and the problems of low accuracy and poor robustness of VAT bills are achieved in the prior art, achieving more efficient correction and higher text recognition accuracy.

CN120013829AInactive Publication Date: 2025-05-16INSPUR SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144305.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has low operating efficiency and poor anti-interference ability in VAT inclination correction, resulting in low accuracy and poor robustness of calibration results, which cannot meet the complex VAT inclination requirements.

Method used

The traditional YOLOv7 network is improved by introducing a depth separable convolution module (DSConv) and a global attention module (GAM), and the improved YOLOv7 network is obtained. Combining the continuity and global characteristics of the bill table line, the VAT bill image is detected in a table line and tilt angle correction.

Benefits of technology

It improves the accuracy of VAT invoice tilt correction and the ability to extract network features, enhances the robustness and timeliness of the algorithm, improves the accuracy of bill text recognition, and promotes the informatization and intelligence of VAT invoice audit management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013829A_ABST
    Figure CN120013829A_ABST
Patent Text Reader

Abstract

The invention discloses a value-added tax bill tilt correction method, system and device and a medium, belongs to the technical field of value-added tax bill image recognition, and aims to solve the technical problem of how to improve the precision of value-added tax bill tilt correction and improve the network feature extraction ability at the same time. The collected value-added tax bill images are made into a data set, and the data set is divided into a training set and a test set; in combination with bill table straight line continuity and global characteristics, a depth separable convolution module and a global attention module are introduced to improve a traditional YOLOv7 network to obtain an improved YOLOv7 network, and table straight lines of a value-added tax bill image are detected through the improved YOLOv7 network; and inputting the training set into the improved YOLOv7 network for training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of value-added tax bill image recognition technology, and in particular to a value-added tax bill tilt correction method, system, equipment and medium. Background Art

[0002] With the continuous progress of society, the use of bills is becoming more and more frequent. Whether it is for commercial capital transactions or capital management, it is inseparable from the support of bills. Therefore, financial personnel inevitably have to spend a lot of energy on the audit and management of bills. In order to realize the intelligent and information-based financial management, many researchers have proposed a series of OCR bill recognition algorithms based on deep learning, cloud computing and other technologies. Bill tilt correction, as an integral part of the preprocessing of bill recognition algorithms, is the cornerstone for improving text positioning accuracy and text recognition accuracy, and has important research significance. In the real VAT bill tilt correction scenario, due to the complexity of the VAT bill face structure, that is, there are a lot of noise interference, text interference and other factors, resulting in traditional tilt correction algorithms such as Hough transform algorithm, Radon transform algorithm, projection transformation algorithm based on perspective transformation, etc., straight line fitting algorithm, directional white run algorithm, etc., in this application scenario, all have the defects of low operating efficiency and poor anti-interference ability, resulting in low accuracy and poor robustness of bill correction results, which cannot meet the needs of complex VAT bill tilt correction.

[0003] With the rise of deep learning, researchers have found that applying convolutional neural networks (CNN) to image tilt correction performs better than traditional algorithms. However, faced with complex ticket structures, they can only improve correction accuracy by stacking basic convolution modules, which undoubtedly leads to deeper and larger parameters in the constructed network model. Not only is the timeliness of model prediction worse, but it also increases the difficulty of training the network model.

[0004] Therefore, how to improve the accuracy of VAT invoice tilt correction and at the same time improve the ability to extract network features is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The technical task of the present invention is to provide a method, system, device and medium for tilt correction of value-added tax bills to solve the problem of how to improve the accuracy of tilt correction of value-added tax bills and at the same time improve the ability to extract network features.

[0006] The technical task of the present invention is achieved in the following manner: a method for correcting the tilt of a value-added tax bill, the method being specifically as follows:

[0007] Collecting VAT invoice images, making the collected VAT invoice images into a data set, and dividing the data set into a training set and a test set;

[0008] Combining the continuity and global characteristics of the straight lines in the bill form, the traditional YOLOv7 network is improved by introducing the deep separable convolution module (DSConv) and the global attention module (GAM) to obtain the improved YOLOv7 network. The improved YOLOv7 network is used to detect the straight lines in the form of the VAT bill image.

[0009] The training set is input into the improved YOLOv7 network for training, and the trained improved YOLOv7 network is verified using the test set to generate a more stable improved YOLOv7 network for table straight line detection of VAT bill images;

[0010] Detect the straight lines in the VAT bills based on the trained improved YOLOv7 network;

[0011] According to the VAT invoice form straight line detection result of the trained improved YOLOv7 network, a rectangular anchor frame of the table straight line is selected, and the classification category, center point coordinates, width, height and confidence of the anchor frame are obtained in the form of (cls, x, y, w, h, conf). The VAT invoice form straight line detection anchor frame is analyzed and calculated, and then the inclination angle of the anchor frame is statistically calculated. The inclination angle of the VAT invoice is corrected according to the statistically calculated anchor frame inclination angle.

[0012] As a preferred method, the collected VAT bill images are made into a data set, and the data set is divided into a training set and a test set as follows:

[0013] Perform binary image preprocessing on the collected VAT bill image, and convert the color VAT bill image into a black and white image;

[0014] Use the labelme annotation tool to select and annotate the black and white VAT bill image, mark the vertex positions of the circumscribed rectangle of the table straight line in the VAT bill, and classify the polygons of the marked table straight line according to the bill inclination angle. The situation is as follows: if the inclination direction of the table straight line is clockwise, it is marked as "Slope0"; if the inclination direction of the table straight line is counterclockwise, it is marked as "Slope1";

[0015] After the labeling is completed, the labelme labeling tool generates a json file containing the labeling data, parses the json file, and generates the corresponding label image and the vertex coordinates of the selected area;

[0016] The processed data set is divided into a training set: test set ratio of 9:1 to generate training sets and test sets for improved YOLOv7 network training.

[0017] Preferably, the global attention module includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearrangement, transforming the input feature map from C×H×W dimensions to W×H×C dimensions; where W represents the width of the feature map; H represents the height of the feature map; and C represents the number of channels of the feature map. A multi-layer perception network (MLP, fully connected network structure) is then used to amplify the correlation of multi-dimensional features; the feature map is then restored to C×H×W dimensions through a reverse three-dimensional rearrangement method; and the output feature map is finally obtained through a sigmoid activation layer.

[0018] A deep separable convolution module is introduced to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule; based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain an improved CBS feature extraction module, thereby improving the YOLOv7 network's ability to extract global features; wherein the improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, thereby optimizing the YOLOv7 network's sensitivity to global table straight line features;

[0019] The design formula of the global attention module is as follows:

[0020]

[0021] Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule.

[0022] More preferably, the depthwise separable convolution module reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution (DW) and pointwise convolution (PW);

[0023] Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map.

[0024] Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

[0025] As a preferred method, the straight line detection anchor frame of the VAT bill form is analyzed and calculated, and then the tilt angle of the anchor frame is statistically calculated. The tilt angle of the VAT bill is corrected according to the statistically calculated anchor frame tilt angle as follows:

[0026] According to the improved YOLOv7 network, the straight line detection anchor frame of the VAT bill form is obtained, and the detection anchor frames with confidence values ​​less than the threshold are filtered out according to the confidence values ​​of the anchor frames to obtain the remaining detection anchor frames;

[0027] Determine whether the tilt direction of the table line is clockwise or counterclockwise, and after determining the tilt direction of the table line, only retain the detection anchor frame corresponding to the tilt direction, and filter out the detection anchor frame of the other tilt direction;

[0028] For the retained detection anchor box, calculate the coordinate values ​​of the lower left corner and the upper right corner of the anchor box rectangle according to the arrays (x, y, w, h) corresponding to the coordinates of the center point, width, and height of the anchor box rectangle, calculate the inclination angle between the straight line connecting the two points and the horizontal direction according to the coordinate values ​​of the lower left corner and the upper right corner of the anchor box rectangle, and use the inclination angle between the straight line connecting the two points of the lower left corner and the upper right corner of the anchor box rectangle and the horizontal direction as the predicted inclination angle of the corresponding anchor box;

[0029] The predicted tilt angles of all retained anchor frames are statistically calculated, the maximum and minimum predicted tilt angles are removed, and the average of the remaining predicted tilt angles is calculated. The calculation results are checked for root mean square error (RMSE), and a root mean square error threshold is set. If the root mean square error is less than the root mean square error threshold, it means that the calculation result meets the correction angle accuracy, and the VAT bill is corrected with the corresponding calculation result; the root mean square error calculation formula is as follows:

[0030]

[0031] Among them, θ RMSE represents the root mean square error; θ mean Represents the average value of the tilt angle θ; N represents the number of detection anchor boxes after filtering, and the maximum and minimum values ​​of the calculated tilt angle are filtered out.

[0032] Preferably, the determination of whether the inclination direction of the table line is clockwise or counterclockwise is as follows:

[0033] Count the detection category cls values ​​of the remaining detection anchor boxes, and determine the relationship between the category with the detection category cls value of Slope0 and the category with the detection category cls value of Slope1:

[0034] If the proportion of the category with the detection category cls value of Slope0 is greater than the proportion of the category with the detection category cls value of Slope1, the inclination direction of the straight line in the table is clockwise;

[0035] If the proportion of the detection category with the cls value of Slope0 is less than the proportion of the detection category with the cls value of Slope1, the inclination direction of the straight line in the table is counterclockwise.

[0036] A VAT bill tilt correction system, which is used to implement the VAT bill tilt correction method as described above, and comprises:

[0037] An image acquisition module, used for acquiring images of value-added tax receipts;

[0038] Correction module: Combining the straight line continuity and global features of the bill form, the traditional YOLOv7 network is improved by introducing the deep separable convolution module (DSConv) and the global attention module (GAM) to obtain an improved YOLOv7 network. Based on the improved YOLOv7 network, the collected VAT bill image is subjected to the straight line inclination angle detection, and the accurate VAT bill image correction angle is obtained through the statistical calculation method of the inclination angle, and the VAT bill image is corrected according to the correction angle;

[0039] The front-end module is used to build the front-end page structure based on the Vue architecture according to business needs;

[0040] The backend module is used to design the backend business logic based on the Django architecture according to business requirements, control the scheduling and execution of the algorithm, save the corrected VAT invoice image to the file server, and return the URL address of the file storage. The frontend echoes the corrected VAT invoice image;

[0041] The deployment module is used to adopt a clustered deployment solution based on Nginx in order to ensure the robustness of the system and effectively improve the stability of the system. The hardware equipment for building the Nginx cluster includes the master proxy server, slave proxy server, static resource server, file storage server, streaming media server and ordinary server.

[0042] Preferably, the global attention module includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearrangement, and transforms the input feature map from C×H×W dimensions to W×H×C dimensions; wherein W represents the width of the feature map; H represents the height of the feature map; C represents the number of channels of the feature map; a multi-layer perception network (MLP, fully connected network structure) is then used to amplify the correlation of multi-dimensional features; the feature map is then restored to C×H×W dimensions by reverse three-dimensional rearrangement; and finally, the output feature map is obtained through a sigmoid activation layer;

[0043] A deep separable convolution module is introduced to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule; based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain an improved CBS feature extraction module, thereby improving the YOLOv7 network's ability to extract global features; wherein the improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, thereby optimizing the YOLOv7 network's sensitivity to global table straight line features;

[0044] The design formula of the global attention module is as follows:

[0045]

[0046] Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule;

[0047] The depthwise separable convolution module reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution (DW) and pointwise convolution (PW);

[0048] Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map.

[0049] Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

[0050] An electronic device, characterized in that it comprises: a memory and at least one processor;

[0051] Wherein, the memory stores a computer program;

[0052] The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the VAT bill tilt correction method as described above.

[0053] A computer-readable storage medium having a computer program stored therein, wherein the computer program can be executed by a processor to implement the VAT invoice tilt correction method as described above.

[0054] The VAT bill tilt correction method, system, device and medium of the present invention have the following advantages:

[0055] (1) The present invention reconstructs the basic convolution unit of the traditional YOLOv7 network by introducing a depthwise separable convolution module (Depthwise SeparableConvolution) and a global attention module GAM (Global Attention Mechanism), expands the network layer of the YOLOv7 network, strengthens the sensitivity of global features, improves the ability to extract network features, and greatly improves the accuracy of VAT bill tilt correction with fewer network model parameters;

[0056] (ii) Based on the global characteristics of the straight lines in the VAT invoice table, the present invention improves the traditional YOLOv7 network by introducing a global attention module (GAM), thereby extending the feature correlation in the global range, improving the ability to extract the features of the straight lines in the table, and realizing effective optimization of the algorithm model for specific scenarios. At the same time, it improves the feature expression ability and execution efficiency of the YOLOv7 network, and has high robustness and timeliness;

[0057] (III) The present invention performs tilt correction on VAT bills with complex face structures, improves the processing capability of the pre-processing stage in the VAT bill text recognition process, improves the accuracy of VAT bill OCR text recognition, and promotes the VAT bill audit and management business to move towards informatization and intelligence;

[0058] (IV) Based on the complex structure of VAT bills and the continuity and globality of bill table lines, the present invention optimizes and reconstructs the traditional YOLOv7 network. On the one hand, in view of the complex structure of VAT bills, the present invention introduces a deep separable convolution module to expand the depth of the network model, effectively reducing the amount of training parameters of the global attention module, improving the algorithm execution efficiency and the feature extraction capability; on the other hand, considering the continuity and globality of bill table lines, the present invention introduces a global attention module to enhance the correlation of straight line features in the global range, thereby improving the accuracy of table straight line detection;

[0059] (V) Based on the table straight line detection results of the improved YOLOv7 network, the present invention proposes a statistical calculation method for the tilt correction angle of the value-added tax bill, adopts a confidence threshold to filter out negative results, and then calculates the weighted average as the tilt correction angle, and verifies the result by calculating the root mean square error, thereby ensuring the accuracy of the tilt correction angle and improving the precision of the correction angle;

[0060] (VI) The present invention effectively combines the front-end module, the back-end module and the correction module, and can be used as a preprocessing tool for value-added tax invoice text recognition. Before performing text recognition on the invoice image, the invoice correction is first performed through the present invention, which can effectively improve the accuracy of the subsequent text recognition algorithm; therefore, the present invention is a supplement and improvement to the value-added tax invoice recognition system. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] The present invention is further described below in conjunction with the accompanying drawings.

[0062] Attached Figure 1 Design architecture diagram for GAM module;

[0063] Attached Figure 2 This is the structure diagram of the depth-separable convolution module;

[0064] Attached Figure 3 This is the structure diagram of the improved GAM spatial attention submodule;

[0065] Attached Figure 4 This is the structure diagram of the optimized CBS feature extraction module;

[0066] Attached Figure 5 This is the training result diagram of the optimized YOLOv7 network;

[0067] Attached Figure 6 This is the straight line detection result diagram of the VAT bill form;

[0068] Attached Figure 7 This is the tilt correction result diagram of the VAT bill;

[0069] Attached Figure 8This is a schematic diagram of the framework of the VAT invoice tilt correction detection system. DETAILED DESCRIPTION

[0070] The VAT invoice tilt correction method, system, device and medium of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.

[0071] Example 1

[0072] This embodiment provides a method for correcting the tilt of a VAT bill, which is specifically as follows:

[0073] S1. Collect VAT bill images, make the collected VAT bill images into a data set, and divide the data set into a training set and a test set;

[0074] S2. Combining the continuity and global features of the straight line of the bill form, the traditional YOLOv7 network is improved by introducing the deep separable convolution module (DSConv) and the global attention module (GAM) to obtain an improved YOLOv7 network, and the straight line of the form of the value-added tax bill image is detected by the improved YOLOv7 network;

[0075] S3, input the training set into the improved YOLOv7 network for training, and use the test set to verify the trained improved YOLOv7 network, so as to generate a more stable improved YOLOv7 network for table straight line detection of VAT bill images;

[0076] S4, detecting the straight lines in the VAT bill based on the trained improved YOLOv7 network;

[0077] S5. According to the VAT invoice form straight line detection result of the trained improved YOLOv7 network, a rectangular anchor frame of the straight line of the form is selected, and the classification category, center point coordinates, width, height and confidence of the anchor frame are obtained in the form of (cls, x, y, w, h, conf). The VAT invoice form straight line detection anchor frame is analyzed and calculated, and the inclination angle of the anchor frame is statistically calculated. The inclination angle of the VAT invoice is corrected according to the statistically calculated anchor frame inclination angle.

[0078] In step S1 of this embodiment, the collected VAT bill images are made into a data set, and the data set is divided into a training set and a test set as follows:

[0079] S101, performing binary image preprocessing on the collected VAT bill image, converting the color VAT bill image into a black and white image;

[0080] S102. Use the labelme annotation tool to select and annotate the black and white VAT bill image, annotate the vertex positions of the circumscribed rectangle of the table straight line in the VAT bill, and classify the polygons of the marked table straight line according to the bill inclination angle. The situation is as follows: if the inclination direction of the table straight line is clockwise, it is marked as "Slope0"; if the inclination direction of the table straight line is counterclockwise, it is marked as "Slope1";

[0081] S103, after the labeling is completed, the labelme labeling tool generates a json file containing the labeling data, parses the json file, and generates a corresponding label image and vertex coordinates of the selected area;

[0082] S104, dividing the processed data set into a training set: test set ratio of 9:1 to generate a training set and a test set for improved YOLOv7 network training.

[0083] The YOLOv7 network is an efficient and accurate target detection network model. It has significant advantages in network structure expansion and feature extraction through the SPP-PANet multi-scale feature fusion algorithm and adaptive convolution module design. However, for the application scenario of VAT bills with complex structure and strong global feature correlation, it is necessary to further improve the model's feature extraction ability and global feature fusion ability.

[0084] In order to solve the problem of the global nature of the VAT bill form, this embodiment introduces a global attention module (GAM module), which can amplify the global dimension interaction characteristics while reducing information dispersion. The GAM module consists of a channel attention submodule and a spatial attention submodule. The design architecture of the global attention module is shown in the attached figure. Figure 1 shown.

[0085] The global attention module in step S2 of this embodiment includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearranging. Figure 1 The permutation structure in the transform transforms the input feature map from C×H×W dimensions to W×H×C dimensions, where W represents the width of the feature map, H represents the height of the feature map, and C represents the number of channels of the feature map. A multi-layer perception network (MLP, fully connected network structure) is then used to amplify the correlation of multi-dimensional features. The feature map is then restored to C×H×W dimensions through a reverse three-dimensional rearrangement method. Finally, the output feature map is obtained through a sigmoid activation layer.

[0086] The spatial attention submodule uses two convolutional layers to fuse spatial information. First, a convolutional layer with a 7×7 convolution kernel is used to reduce the dimension according to the dimension (number of channels C) reduction ratio r, and the reduction ratio coefficient set in the present invention is 4; then, a convolutional layer with a 7×7 convolution kernel is used to restore the previous dimension; finally, the output feature map is obtained through a sigmoid activation layer.

[0087] The design formula of the global attention module is as follows:

[0088]

[0089] Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule.

[0090] By introducing the global attention module, the global dimension information and spatial information are effectively integrated, which improves the network model's ability to extract global features. However, the global attention module introduces two convolutional layers with 7×7 convolution kernels. At the same time, considering that the YOLOv7 network upgrades the feature map to a larger number of channels during feature extraction, the introduction of the global attention module alone will increase the number of parameters in the network model and reduce its operating efficiency.

[0091] In order to reduce the negative effect of the sudden increase in the number of YOLOv7 network parameters after the introduction of the global attention module, this embodiment introduces a depthwise separable convolution module to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule.

[0092] As attached Figure 2 As shown, the depthwise separable convolution module in step S2 of this embodiment reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution (DW) and pointwise convolution (PW);

[0093] Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map.

[0094] Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

[0095] In order to illustrate the effects of depthwise separable convolution and traditional convolution on the parameters of the YOLOv7 network, this embodiment takes an input feature map of 32×256×256 dimensions as an example, uses a 7×7 convolution kernel for convolution, and outputs an output feature map of 64×256×256 dimensions. Substitute it into the description of depthwise separable convolution, C=32; H=256; W=256; M=64; then, use depthwise separable convolution and traditional convolution to calculate the parameters of the YOLOv7 network, and obtain the traditional convolution parameter N. Conv and the number of depth-wise separable convolution parameters N DSConv , the calculation formula is as follows:

[0096] N Conv =7×7×32×64=100352;

[0097] N DSConv =N DW +N PW =7×7×32+64×1×1=1632.

[0098] Therefore, by introducing a depthwise separable convolutional unit to replace the traditional convolutional unit of the spatial attention submodule in the global attention module, the problem of a sharp increase in model parameters and reduced operating efficiency caused by the introduction of the global attention module can be effectively solved. The improved GAM spatial attention submodule is shown in the attached figure. Figure 3 shown.

[0099] Based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain the improved CBS feature extraction module, which improves the YOLOv7 network's ability to extract global features. The optimized CBS feature extraction module is shown in the attached figure. Figure 4 As shown in the figure, the traditional CBS feature extraction module is composed of a convolutional layer (Conv), a normalization layer (Batch Normalization), and a SiLu activation function, and is named by the initials of these three parts. The improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, which optimizes the sensitivity of the YOLOv7 network to global table straight line features.

[0100] In step S3 of this embodiment, the training set is input into the improved YOLOv7 network for training. Specifically, the hardware environment for network training is Ubuntu operating system and NVIDIA GeForce GTX1080Ti graphics card; the software environment is Python, and the deep learning part is implemented by PyTorch framework programming. During the model training process, the initialization learning rate lr0 is set to 0.01, the maximum iteration cycle Epoch is 1000, and the training batch batch_size is 16. After the training is completed, the model training results are evaluated by the convergence of the loss function and the precision index. The training results obtained are shown in the attached figure. Figure 5 As shown. Among them, the precision rate refers to the ratio of the number of correctly predicted positive samples to the number of predicted positive samples, which indicates the probability of actually being positive samples among all samples predicted to be positive. The calculation formula is as follows:

[0101]

[0102] Among them, TP (True Positives) represents the number of samples correctly classified as positive, that is, the number of samples whose actual category and predicted category are both positive; FP (False Positives) represents the number of samples incorrectly classified as positive, that is, the number of samples whose actual category is negative but the predicted category is positive.

[0103] This embodiment optimizes the training parameters of the improved YOLOv7 network to obtain the optimal improved YOLOv7 network training weights, and applies the trained improved YOLOv7 network to the straight line detection of the VAT bill form. The detection results are shown in the attached figure. Figure 6 As shown in the figure. The experimental results show that the improved GAM attention module can be used to globally optimize the YOLOv7 network. Even if the VAT bill form line is long, it can accurately locate the form line position and accurately detect the VAT bill form straight line. At the same time, it can accurately classify the inclination angle as clockwise or counterclockwise. Therefore, the improved YOLOv7 network is more suitable for the detection of VAT bill form straight lines, and has higher accuracy and robustness.

[0104] In step S5 of this embodiment, the straight line detection anchor frame of the VAT bill form is analyzed and calculated, and then the tilt angle of the anchor frame is statistically calculated. The tilt angle of the VAT bill is corrected according to the statistically calculated anchor frame tilt angle as follows:

[0105] S501, obtaining a straight line detection anchor frame of the VAT invoice form according to the improved YOLOv7 network, the confidence threshold set in this embodiment is 0.8; filtering out the detection anchor frames whose confidence is less than the threshold according to the confidence conf value of the anchor frame, and obtaining the remaining detection anchor frames;

[0106] S502, determining whether the tilt direction of the table straight line is clockwise or counterclockwise, and after determining the tilt direction of the table straight line, retaining only the detection anchor frame corresponding to the tilt direction, and filtering out the detection anchor frame of the other tilt direction; wherein determining whether the tilt direction of the table straight line is clockwise or counterclockwise is specifically as follows:

[0107] Count the detection category cls values ​​of the remaining detection anchor boxes, and determine the relationship between the category with the detection category cls value of Slope0 and the category with the detection category cls value of Slope1:

[0108] If the proportion of the category with the detection category cls value of Slope0 is greater than the proportion of the category with the detection category cls value of Slope1, the inclination direction of the straight line in the table is clockwise;

[0109] If the proportion of the category with the detection category cls value of Slope0 is less than the proportion of the category with the detection category cls value of Slope1, the inclination direction of the straight line in the table is counterclockwise;

[0110] S503, for the retained detection anchor frame, calculate the coordinate values ​​of the lower left corner and the upper right corner of the anchor rectangle according to the arrays (x, y, w, h) corresponding to the coordinates of the center point, width and height of the anchor rectangle, calculate the inclination angle between the straight line connecting the two points and the horizontal direction according to the coordinate values ​​of the lower left corner and the upper right corner of the anchor rectangle, and use the inclination angle between the straight line connecting the two points of the lower left corner and the upper right corner of the anchor rectangle and the horizontal direction as the predicted inclination angle of the corresponding anchor frame;

[0111] S504. Perform statistical calculations on the predicted tilt angles of all retained anchor frames, remove the maximum and minimum values ​​of the predicted tilt angles, calculate the average of the remaining predicted tilt angles, and perform a root mean square error (RMSE) check on the calculation result. Considering that the tilt angle value is generally small and is greatly affected by the error, the root mean square error threshold is set to 5% of the average value in this embodiment. If the root mean square error is less than 5% of the average value, it is considered that the calculation result meets the correction angle accuracy, and finally the value-added tax invoice is corrected with the calculation result.

[0112] The experimental results of the improved YOLOv7 network test in this embodiment (corresponding to the attached Figure 6 ) as an example. The above calculation process is elaborated in detail:

[0113] Based on the experimental results of the improved YOLOv7 network test in this embodiment, the predicted anchor frame data of the optimized YOLOv7 network is obtained, and the anchor frame data is statistically calculated, as shown in Table 1.

[0114] Table 1. Statistical calculation table of predicted anchor boxes of YOLOv7 network

[0115]

[0116]

[0117] According to the inclination angle θ calculated in Table 1, the maximum value (1.987°) and the minimum value (1.695°) of the inclination angle θ are removed, and then the average value of the remaining inclination angle θ is calculated to be 1.807°. In order to ensure the calculation accuracy of the inclination angle θ, the root mean square error is used to verify the inclination angle θ. The calculation formula of the root mean square error is as follows:

[0118]

[0119] Among them, θ RMSE represents the root mean square error; θ mean represents the average value of the tilt angle θ; N represents the number of detected anchor frames after filtering, and the maximum and minimum values ​​of the calculated tilt angle are filtered out, corresponding to the number of anchor frames after filtering in Table 1, which is 4; substituting the data in Table 1, the root mean square error value θ can be obtained. RMSE The result is 0.0487°, which is less than 5% of the average value and is within the reasonable error range. Therefore, the statistical calculation shows that the inclination angle of the VAT bill is 1.807° counterclockwise (corresponding to category "Slope1").

[0120] Through this embodiment, the tilt correction result of the value-added tax bill can be obtained, as shown in the attached figure. Figure 7 shown.

[0121] Embodiment 2:

[0122] As attached Figure 8 The present embodiment provides a VAT bill tilt correction system, which is used to implement the VAT bill tilt correction method in Example 1, and the system includes:

[0123] An image acquisition module, used for acquiring images of value-added tax receipts;

[0124] Correction module: Combining the straight line continuity and global features of the bill form, the traditional YOLOv7 network is improved by introducing the deep separable convolution module (DSConv) and the global attention module (GAM) to obtain an improved YOLOv7 network. Based on the improved YOLOv7 network, the collected VAT bill image is subjected to the straight line inclination angle detection, and the accurate VAT bill image correction angle is obtained through the statistical calculation method of the inclination angle, and the VAT bill image is corrected according to the correction angle;

[0125] The front-end module is used to build the front-end page structure based on the Vue architecture according to business needs;

[0126] The backend module is used to design the backend business logic based on the Django architecture according to business requirements, control the scheduling and execution of the algorithm, save the corrected VAT invoice image to the file server, and return the URL address of the file storage. The frontend echoes the corrected VAT invoice image;

[0127] The deployment module is used to adopt a clustered deployment solution based on Nginx in order to ensure the robustness of the system and effectively improve the stability of the system. The hardware equipment for building the Nginx cluster includes the master proxy server, slave proxy server, static resource server, file storage server, streaming media server and ordinary server.

[0128] The global attention module in this embodiment includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearrangement, and transforms the input feature map from C×H×W dimensions to W×H×C dimensions; wherein W represents the width of the feature map; H represents the height of the feature map; and C represents the number of channels of the feature map. A multi-layer perception network (MLP, fully connected network structure) is then used to amplify the correlation of multi-dimensional features; the feature map is then restored to C×H×W dimensions by reverse three-dimensional rearrangement; and finally, the output feature map is obtained through a sigmoid activation layer.

[0129] A deep separable convolution module is introduced to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule; based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain an improved CBS feature extraction module, thereby improving the YOLOv7 network's ability to extract global features; wherein the improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, thereby optimizing the YOLOv7 network's sensitivity to global table straight line features;

[0130] The design formula of the global attention module is as follows:

[0131]

[0132] Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule;

[0133] The depthwise separable convolution module reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution (DW) and pointwise convolution (PW);

[0134] Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map.

[0135] Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

[0136] The working process of the system is as follows:

[0137] Step 1: Collect VAT receipt images;

[0138] Step 2: Build a front-end page based on the Vue architecture; upload the VAT invoice image to the back-end through the file upload component, and echo the VAT invoice image corrected by the algorithm;

[0139] Step 3: Build the system backend based on the Django architecture, embed the VAT bill tilt correction algorithm and hardware device control, and connect the front-end, back-end, database and other data access services in series;

[0140] Step 4: System testing: According to the functional requirements of each functional module, test the execution of the system and modify the existing bugs and derivative problems in the system;

[0141] Step 6. System deployment: In order to improve the stability and robustness of the system, the system deployment adopts a clustered deployment solution based on Nginx. The deployment architecture is mainly composed of a master proxy server, a slave proxy server, a static resource server, a file storage server, a streaming media server, and a back-end ordinary server. Among them, the file storage server saves the VAT invoice after the algorithm model tilt correction.

[0142] Based on computer vision technology and deep learning technology as the algorithm, combined with software development technology, a VAT invoice tilt correction detection system was developed. On the one hand, the algorithm was implemented and the stability and reliability of the algorithm were tested; on the other hand, the system can be used as a preprocessing tool for VAT invoice text recognition. Before text recognition is performed on the invoice image, the invoice is corrected through this system, which can effectively improve the accuracy of the subsequent text recognition algorithm; at the same time, the system is conducive to improving the informatization and intelligence level of VAT invoice reimbursement management business, and has broad application prospects.

[0143] Embodiment 3:

[0144] This embodiment also provides an electronic device, including: a memory and a processor;

[0145] Wherein, the memory stores computer-executable instructions;

[0146] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the value-added tax bill tilt correction method in any embodiment of the present invention.

[0147] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0148] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0149] Embodiment 4:

[0150] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the VAT bill tilt correction method in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0151] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0152] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0153] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0154] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for correcting the tilt of a value-added tax bill, characterized in that: The method is as follows: Collecting VAT invoice images, making the collected VAT invoice images into a data set, and dividing the data set into a training set and a test set; Combining the continuity and global characteristics of the straight lines in the bill form, the traditional YOLOv7 network is improved by introducing a deep separable convolution module and a global attention module to obtain an improved YOLOv7 network. The improved YOLOv7 network is used to detect the straight lines in the form of the VAT bill image. The training set is input into the improved YOLOv7 network for training, and the trained improved YOLOv7 network is verified using the test set to generate a stable improved YOLOv7 network for table straight line detection of VAT bill images; Detect the straight lines in the VAT bills based on the trained improved YOLOv7 network; According to the VAT invoice form straight line detection result of the trained improved YOLOv7 network, a rectangular anchor frame of the table straight line is selected, and the classification category, center point coordinates, width, height and confidence of the anchor frame are obtained in the form of (cls, x, y, w, h, conf). The VAT invoice form straight line detection anchor frame is analyzed and calculated, and then the inclination angle of the anchor frame is statistically calculated. The inclination angle of the VAT invoice is corrected according to the statistically calculated anchor frame inclination angle.

2. The method for correcting the inclination of a value-added tax bill according to claim 1, characterized in that: The collected VAT invoice images are made into a data set, and the data set is divided into a training set and a test set as follows: Perform binary image preprocessing on the collected VAT bill image, and convert the color VAT bill image into a black and white image; Use the labelme annotation tool to select and annotate the black and white VAT bill image, mark the vertex positions of the circumscribed rectangle of the table straight line in the VAT bill, and classify the polygons of the marked table straight line according to the bill inclination angle. The situation is as follows: if the inclination direction of the table straight line is clockwise, it is marked as "Slope0"; if the inclination direction of the table straight line is counterclockwise, it is marked as "Slope1"; After the labeling is completed, the labelme labeling tool generates a json file containing the labeling data, parses the json file, and generates the corresponding label image and the vertex coordinates of the selected area; The processed data set is divided into a training set: test set ratio of 9:1 to generate training sets and test sets for improved YOLOv7 network training.

3. The method for correcting the inclination of a value-added tax bill according to claim 1 or 2, characterized in that: The global attention module includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearrangement, transforming the input feature map from C×H×W dimensions to W×H×C dimensions; where W represents the width of the feature map; H represents the height of the feature map; and C represents the number of channels of the feature map. A multi-layer perception network is then used to amplify the correlation of multi-dimensional features. The feature map is then restored to C×H×W dimensions through a reverse three-dimensional rearrangement method. Finally, the output feature map is obtained through a sigmoid activation layer. A deep separable convolution module is introduced to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule; based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain an improved CBS feature extraction module, thereby improving the YOLOv7 network's ability to extract global features; wherein the improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, thereby optimizing the YOLOv7 network's sensitivity to global table straight line features; The design formula of the global attention module is as follows: Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule.

4. The method for correcting the inclination of a VAT bill according to claim 3, characterized in that: The depthwise separable convolution module reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution and pointwise convolution; Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map. Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

5. The method for correcting the inclination of a value-added tax bill according to claim 1, characterized in that: The straight line detection anchor frame of the VAT bill form is analyzed and calculated, and then the tilt angle of the anchor frame is statistically calculated. The tilt angle of the VAT bill is corrected according to the statistically calculated anchor frame tilt angle as follows: According to the improved YOLOv7 network, the straight line detection anchor frame of the VAT bill form is obtained, and the detection anchor frames with confidence values ​​less than the threshold are filtered out according to the confidence values ​​of the anchor frames to obtain the remaining detection anchor frames; Determine whether the tilt direction of the table line is clockwise or counterclockwise, and after determining the tilt direction of the table line, only retain the detection anchor frame corresponding to the tilt direction, and filter out the detection anchor frame of the other tilt direction; For the retained detection anchor box, calculate the coordinate values ​​of the lower left corner and the upper right corner of the anchor box rectangle according to the arrays (x, y, w, h) corresponding to the coordinates of the center point, width, and height of the anchor box rectangle, calculate the inclination angle between the straight line connecting the two points and the horizontal direction according to the coordinate values ​​of the lower left corner and the upper right corner of the anchor box rectangle, and use the inclination angle between the straight line connecting the two points of the lower left corner and the upper right corner of the anchor box rectangle and the horizontal direction as the predicted inclination angle of the corresponding anchor box; Perform statistical calculations on the predicted tilt angles of all retained anchor frames, remove the maximum and minimum values ​​of the predicted tilt angles, and then calculate the average value of the remaining predicted tilt angles. Perform root mean square error verification on the calculation results, and set a root mean square error threshold. If the root mean square error is less than the root mean square error threshold, it means that the calculation result meets the correction angle accuracy, and the value-added tax bill is corrected with the corresponding calculation result; the root mean square error calculation formula is as follows: Among them, θ RMSE represents the root mean square error; θ mean Represents the average value of the tilt angle θ; N represents the number of detection anchor boxes after filtering, and the maximum and minimum values ​​of the calculated tilt angle are filtered out.

6. The method for correcting the inclination of a value-added tax bill according to claim 5, characterized in that: Determine whether the tilt direction of the table line is clockwise or counterclockwise as follows: Count the detection category cls values ​​of the remaining detection anchor boxes, and determine the relationship between the category with the detection category cls value of Slope0 and the category with the detection category cls value of Slope1: If the proportion of the category with the detection category cls value of Slope0 is greater than the proportion of the category with the detection category cls value of Slope1, the inclination direction of the straight line in the table is clockwise; If the proportion of the detection category with the cls value of Slope0 is less than the proportion of the detection category with the cls value of Slope1, the inclination direction of the straight line in the table is counterclockwise.

7. A VAT bill tilt correction system, characterized in that: The system is used to implement the VAT bill tilt correction method as described in any one of claims 1 to 6, and the system includes: An image acquisition module, used for acquiring images of value-added tax receipts; Correction module: Combining the straight line continuity and global features of the bill form, the traditional YOLOv7 network is improved by introducing a deep separable convolution module and a global attention module to obtain an improved YOLOv7 network. Based on the improved YOLOv7 network, the collected VAT bill image is subjected to the straight line inclination angle detection, and the accurate VAT bill image correction angle is obtained through the statistical calculation method of the inclination angle, and the VAT bill image is corrected according to the correction angle; The front-end module is used to build the front-end page structure based on the Vue architecture according to business needs; The backend module is used to design the backend business logic based on the Django architecture according to business requirements, control the scheduling and execution of the algorithm, save the corrected VAT invoice image to the file server, and return the URL address of the file storage. The frontend echoes the corrected VAT invoice image; The deployment module is used to adopt a clustered deployment solution based on Nginx. The hardware equipment for building an Nginx cluster includes a master proxy server, a slave proxy server, a static resource server, a file storage server, a streaming media server, and a common server.

8. The VAT bill tilt correction system according to claim 7, characterized in that: The global attention module includes a channel attention submodule and a spatial attention submodule. The channel attention submodule retains information in three dimensions by rearrangement, transforming the input feature map from C×H×W dimensions to W×H×C dimensions; where W represents the width of the feature map; H represents the height of the feature map; and C represents the number of channels of the feature map. A multi-layer perception network is then used to amplify the correlation of multi-dimensional features. The feature map is then restored to C×H×W dimensions through a reverse three-dimensional rearrangement method. Finally, the output feature map is obtained through a sigmoid activation layer. A deep separable convolution module is introduced to improve the spatial attention submodule of the global attention module to obtain an improved spatial attention submodule; based on the improved spatial attention submodule, the CBS feature extraction module of the YOLOv7 network is optimized to obtain an improved CBS feature extraction module, thereby improving the YOLOv7 network's ability to extract global features; wherein the improved CBS feature extraction module includes an improved spatial attention submodule, a normalization layer, and a SiLu activation function, thereby optimizing the YOLOv7 network's sensitivity to global table straight line features; The design formula of the global attention module is as follows: Among them, F1 represents the input feature map; represents the convolution operation; M C represents the operation process of the channel attention submodule; F2 represents the feature map output by the channel attention submodule; M S represents the operation process of the improved spatial attention submodule; F3 represents the feature map output by the improved spatial attention submodule; The depthwise separable convolution module reconstructs the spatial attention submodule of the global attention module by combining depthwise convolution and pointwise convolution; Among them, deep convolution is used for feature extraction. Convolution is performed in a way that one convolution kernel corresponds to one feature map channel. The convolution kernel size is set to 7×7. One channel is convolved by only one convolution kernel, so that the output feature map of the deep convolution has the same number of channels as the input feature map. Point-by-point convolution is applied to multi-channel feature fusion. The output feature map of the deep convolution is convolved with M 1×1 convolution kernels. The dimension of the output feature map after point-by-point convolution is M, and the feature map of C×H×W dimensions is transformed into a feature map of M×H×W dimensions through deep convolution transformation.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the VAT bill tilt correction method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the VAT bill tilt correction method as described in any one of claims 1 to 6.