An invoice text information recognition method and system
By constructing a dynamic convolution YOLOv8 model and TPS STN, combined with local hybrid strategy, the text detection and recognition problem of taxi invoice images is solved, and more efficient text information extraction and recognition is achieved.
Patent Information
- Application Number
- CN202510577886.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When processing taxi invoice images, existing deep learning models have problems such as insufficient preprocessing of image preprocessing, limited text detection accuracy and accumulated text recognition errors, especially in complex scenarios, it is difficult to accurately extract text areas and recognize characters.
A dynamic convolution YOLOv8 model based on ParameterNet is constructed, combined with TPS's STN and SVTR, and through dynamic convolution and local mixing strategies, the model's feature extraction and text recognition capabilities are enhanced, and text areas of different shapes and sizes are adapted, and local feature perception capabilities are improved through local mixing.
It improves the accuracy and robustness of taxi invoice text detection, reduces the complexity of model calculation, enhances the adaptability to complex backgrounds and geometric deformations, and improves text recognition efficiency.
Smart Images

Figure CN120088810B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of invoice recognition, and in particular, to a method and system for recognizing invoice text information. Background Art
[0002] With the development of information technology, financial reimbursement has gradually stepped onto the path of informatization. Traditional manual entry of bill information will gradually be replaced by intelligent recognition. The digitization of bills and the text detection and recognition of electronic bills directly determine the feasibility of intelligent financial reimbursement. As one of many bills, taxi invoices have their own uniqueness and complexity in image text detection and recognition due to the particularity of their layout format.
[0003] In real life, the text layout of original taxi invoices is diverse. Some text lines have inconsistent line spacing, some text blocks are irregularly distributed, some texts are dense, resulting in a large amount of text information concentrated in a limited area, and there is also uncertainty in the quality of invoice images, such as blurred images, uneven illumination, and stains. These factors all affect the accuracy and robustness of text detection. Traditional OCR technology relies on manual feature extraction and template matching, so it is difficult to handle problems such as text density, blurred fonts, skewed deformation, and complex background interference in taxi invoice images.
[0004] Some existing end-to-end models based on deep learning (such as YOLOv8, DBNet, CRNN, etc.) have gradually become the mainstream solutions for invoice text recognition. However, there are still the following key defects when dealing with complex scenarios such as taxi invoices:
[0005] 1. Insufficient robustness in image preprocessing. Due to factors such as shooting environment and storage conditions, taxi invoices often have severe interference. The problem of uneven illumination makes it difficult to accurately extract the text area, and there are folding and stain interferences.
[0006] 2. Limited accuracy in text detection. In the dense area of the table lines of taxi invoices, due to the extremely small line spacing of text lines, the detection method based on anchor boxes is prone to overlapping box merging errors. For key information such as invoice codes, due to the small character size, lightweight models (such as YOLOv5s) have insufficient receptive fields, resulting in significant positioning offsets. And for some slanted and curved texts, the mainstream horizontal rectangle box detection cannot adapt and must rely on an additional angle regression module (such as RRPN), thus increasing the computational cost.
[0007] 3. Cumulative errors in text recognition. The text recognition stage is limited by font diversity, low resolution, and semantic relevance. Complex fonts in taxi invoices, such as similar characters (such as "3" and "8"), are prone to misjudgment. Low-quality images such as blurred and faded texts result in the failure of convolutional feature extraction, and recognition is prone to errors. Summary of the Invention
[0008] The present invention aims to at least improve one of the technical problems existing in the prior art. To this end, the present invention proposes an invoice text information recognition method and system.
[0009] The technical solution of the present invention is as follows:
[0010] An invoice text information recognition method, which includes:
[0011] Obtain an invoice image as a detection data set, and annotate the key text in the invoice image to obtain training samples;
[0012] Construct a YOLOv8 model based on the dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. Introduce the dynamic convolution in ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model;
[0013] Introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and train the YOLOv8 model according to the training samples to output the text area of the training samples;
[0014] Construct an invoice text information recognition data set according to the text area;
[0015] Construct a text recognition model. The text recognition model includes a spatial transformation network (STN) based on thin plate spline interpolation (TPS) and a scene text recognition network (SVTR) based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; Correct the shape of the recognition data set according to the spatial transformation network, and enhance the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR;
[0016] Input the invoice text information recognition data set into the text recognition model to output structured invoice text information.
[0017] In a possible technical solution, further, the spatial transformation network includes a localization network, a grid generator, and a sampler, where the localization network is a convolutional neural network;
[0018] Specifically, correcting the shape of the recognition data set according to the spatial transformation network includes:
[0019] The positioning network is used to receive the input image in the invoice text information recognition dataset for predicting the positions of a set of key control points in the input image, where the key control points represent the areas in the image that need to be deformed;
[0020] The grid generator generates a deformation grid according to the positions of the key control points by using the thin plate spline interpolation algorithm, and defines the new positions of each pixel in the input image in the output corrected image according to the deformation grid;
[0021] The sampler samples the input image according to the deformation grid to generate an output corrected image, making the text arrangement more horizontal and neat.
[0022] In a possible technical solution, further, enhancing the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR specifically includes:
[0023] Receiving the output corrected image to the image block embedding module for division to obtain image blocks of a fixed size;
[0024] Performing convolution mapping according to the image blocks of the fixed size to obtain the initial feature sequence of the output corrected image;
[0025] Passing the initial feature sequence through the first mixing module and the second mixing module to generate a compressed feature sequence. Both the first mixing module and the second mixing module contain multiple local mixings. The first mixing module and the second mixing module mainly combine the local window self-attention mechanism and the feed-forward network to realize local feature modeling and fusion operations, for gradually compressing the length of the feature sequence, improving the calculation efficiency and enhancing the semantic abstraction ability;
[0026] Passing the compressed feature sequence through the third mixing module to generate the final feature sequence. The third mixing module includes one local mixing and five global mixings;
[0027] Passing the final feature sequence through a fully connected layer for character classification to output the recognition result.
[0028] In a possible technical solution, further, the dynamic convolution network based on ParameterNet is generated by combining the convolution kernels and weight generation networks of multiple dynamic experts (DEs), through multiple predefined convolution kernels and the dynamic coefficients corresponding to the convolution kernels through weighted combination;
[0029] The output feature map of the YOLOv8 model based on the dynamic convolution network The calculation formula is:
[0030]
[0031] Among them, represents the input feature map, and its dimensions are represented as , represents the number of input channels, H represents the input height, and W represents the input width; represents the output feature map, which is the output result of the dynamic convolution operation, and its dimensions are represented as , represents the number of output channels, represents the output height, , represents the i-th dynamic expert, and M represents the number of dynamic experts;
[0032] Dynamic coefficient is generated based on the input features through global average pooling and a multi-layer perceptron, and the formula for generating is:
[0033] Pool represents global average pooling, and MLP represents a multi-layer perceptron. Dynamic convolution can dynamically generate convolution kernel parameters according to the input features, so as to better adapt to text regions of different shapes and sizes, and enhance the robustness of the model to deformation and noise. At the same time, the parameterization mechanism provided by ParameterNet can effectively control the computational complexity of dynamic convolution, so that while maintaining high performance, it will not significantly increase the computational burden of the model, and thus is more suitable for resource-constrained practical application scenarios.
[0034] In a possible technical solution, further, during the annotation process, rectangular bounding boxes are used to frame the key text of each invoice image to obtain the corresponding target label file, thereby obtaining training samples; among them, when framing, the framing range can completely cover the key text, while minimizing the interference of irrelevant backgrounds, so as to improve the accuracy and robustness of subsequent model training.
[0035] In a possible technical solution, further, during the feature extraction stage of the backbone network, it includes:
[0036] The training samples are passed through multiple downsampling convolution operations to extract multi-scale features. Among them, the convolution kernel size of the downsampling convolution is 3, and the stride is 2, which is used to reduce the feature map size and increase the number of channels;
[0037] Intersperse dynamic convolutions (C2f-DynamicConv) among multiple downsampling convolutions in the backbone network to extract features at different scales, and finally enter the SPPF module in the YOLOv8 model to output the feature results, completing the feature extraction process of the backbone network. By interspersing multiple C2f-DynamicConv in the backbone network, richer features can be extracted more effectively, enabling it to better adapt to different inputs and improving the efficiency of the model. As the network deepens, the size of the feature map gradually decreases while the number of channels gradually increases, and finally enters the SPPF module. The SPPF module uses a multi-scale max-pooling structure to enhance the receptive field and integrate multi-scale context information without changing the size of the feature map, improving the robustness to different target sizes.
[0038] In one possible technical solution, further, in the neck network feature fusion stage, it includes:
[0039] Perform upsampling convolution on the feature results output by the backbone network, and perform a splicing operation with the features at different scales output by the dynamic convolution (C2f-DynamicConv) in the backbone network to obtain the spliced scale features;
[0040] Fuse the different scale features after splicing according to the dynamic convolution (C2f-DynamicConv) to form a rich multi-scale semantic information flow.
[0041] According to the invoice text information recognition method of the present invention, by dynamically generating the convolution kernel weights, the optimized YOLOv8 model can adaptively adjust the convolution operation according to the input, thereby enhancing the expression ability of the model and improving the text detection ability. In addition, by replacing the ordinary convolution in the YOLOv8 detection head with partial convolution (PConv), the amount of computation and the number of parameters are significantly reduced, thereby reducing the complexity of the model and improving the text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction, and utilizes the effectiveness of STN in geometric correction to effectively alleviate geometric deformations such as text distortion and tilt in the input image, thereby improving the text recognition efficiency. By increasing the proportion of the local mixture in the SVTR to enhance the local feature perception ability and text information capture ability of the SVTR, it is more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, and thus can more effectively identify various structured fields on the invoice.
[0042] An invoice text information recognition system, wherein the invoice text information is recognized by using the above recognition method, and the system includes:
[0043] An acquisition module, configured to acquire an invoice image as a detection data set, and label the key text in the invoice image to obtain training samples;
[0044] The first construction module is used to construct a YOLOv8 model with dynamic convolution based on ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. The dynamic convolution in ParameterNet is introduced into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model;
[0045] The training output module is used to introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model's detection head, and train the YOLOv8 model according to the training samples to output the text regions of the training samples;
[0046] The second construction module is used to construct an invoice text information recognition dataset according to the text regions;
[0047] The third construction module is used to construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; the shape of the recognition dataset is corrected according to the spatial transformation network, and the local feature perception ability and text information capture ability of the SVTR are enhanced by increasing the proportion of the number of local mixing in the SVTR;
[0048] The output module is used to input the invoice text information recognition dataset into the text recognition model to output structured invoice text information.
[0049] A computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the invoice text information recognition method as described above.
[0050] A computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer is made to execute the invoice text information recognition method as described above.
[0051] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a flowchart of an invoice text information recognition method according to an embodiment of the present invention;
[0054] Figure 2 is an architecture diagram of a text recognition model of an invoice text information recognition method according to an embodiment of the present invention;
[0055] Figure 3 is a schematic diagram of an invoice text information recognition system according to an embodiment of the present invention. Detailed implementation manners
[0056] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the accompanying drawings are exemplary. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0059] The terms "first", "second", "third", etc. in the description and claims of the present application and the accompanying drawings are used to distinguish different objects and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, steps or units not listed are further included, or optionally, other steps or units inherent to these processes, methods, products or devices are further included.
[0060] Only the parts relevant to the present application are shown in the accompanying drawings, rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and so on.
[0061] The terms "component", "module", "system", "unit", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media storing various data structures. A unit can communicate, for example, through signals with other systems through local and / or remote processes according to signals having one or more data packets (such as data from a second unit interacting with a local system, a distributed system, and / or a network. For example, the Internet interacting with other systems through signals).
[0062] Embodiment 1
[0063] This embodiment provides an invoice text information recognition method, which includes:
[0064] S100, obtaining an invoice image as a detection data set, and annotating key texts in the invoice image to obtain training samples;
[0065] S200, constructing a YOLOv8 model based on dynamic convolution of ParameterNet, the YOLOv8 model includes a backbone network, a neck network, and multiple detection heads, introducing the dynamic convolution in ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model;
[0066] S300, introducing a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, improve the inference speed and resource utilization efficiency of the model, improve the text detection ability, and train the YOLOv8 model according to the training samples to output the text regions of the training samples;
[0067] S400, constructing an invoice text information recognition data set according to the text regions;
[0068] S500. Build a text recognition model, where the text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on the single vision model includes local mixing and global mixing. Correct the shape of the recognition dataset according to the spatial transformation network, and enhance the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR.
[0069] S600. Input the invoice text information recognition dataset into the text recognition model to output structured invoice text information.
[0070] It should be noted that in S100, during the annotation process, a rectangular bounding box is used to select the key text of each invoice image to obtain the corresponding target label file, thereby obtaining training samples. Among them, when selecting, ensure that the selection range can completely cover the key text, and at the same time minimize the interference of irrelevant backgrounds to improve the accuracy and robustness of subsequent model training. After annotation, each box corresponds to a target label, and using the LabelImg annotation tool will automatically generate the corresponding label file. The label file stores the annotation information in text format, with each line corresponding to a target label. The first number represents the target category, and the following four numbers represent the position information of the target bounding box, including the center coordinates (x, y) of the bounding box and the width and height (w, h). These values are all normalized between the width and height of the image, and the value range is [0,1], which is convenient for subsequent deep learning models to read and use.
[0071] It should be noted that in S200, the dynamic convolution network based on ParameterNet is a network that generates convolution kernels and weights by combining multiple Dynamic Experts (DEs). It can dynamically adjust the convolution kernel weights to adapt to the characteristics of the input data, effectively capture the features of small text areas on the invoice, and at the same time adapt to various font and layout forms. In addition, it can effectively suppress the interference of complex backgrounds such as invoice stamps, lines, or patterns, focus on the text area, thereby improving the robustness and accuracy of detection. By enhancing the model's adaptability to geometric changes and detail features, dynamic convolution can significantly improve the accuracy and generalization performance of taxi invoice text detection.
[0072] In the improved YOLOv8 model of this embodiment, for each input feature map X, the model no longer uses a single fixed convolution kernel, but is generated by the weighted combination of multiple predefined convolution kernels and the dynamic coefficients corresponding to the convolution kernels.
[0073] The output feature map of the YOLOv8 model based on the dynamic convolution network The calculation formula is:
[0074]
[0075] Wherein, represents the input feature map, and its dimension is expressed as , represents the number of input channels, H represents the input height, and W represents the input width; represents the output feature map, which is the output result of the dynamic convolution operation, and its dimension is expressed as , represents the number of output channels, represents the output height, , represents the i-th dynamic expert, and M represents the number of dynamic experts;
[0076] Dynamic coefficient is generated based on the input features through global average pooling and a multi-layer perceptron, and the formula for generating is:
[0077] Pool represents global average pooling, and MLP represents a multi-layer perceptron. Dynamic convolution can dynamically generate convolution kernel parameters according to input features, so as to better adapt to text regions of different shapes and sizes, and enhance the robustness of the model to deformation and noise. At the same time, the parameterization mechanism provided by ParameterNet can effectively control the computational complexity of dynamic convolution, so that while maintaining high performance, it will not significantly increase the computational burden of the model, and thus is more suitable for resource-constrained practical application scenarios.
[0078] It should be noted that in this embodiment, in the backbone network feature extraction stage, it includes:
[0079] Training samples are used to extract multi-scale features through multiple downsampling convolution operations, where the convolution kernel size of the downsampling convolution is 3 and the stride is 2, which is used to reduce the feature map size and increase the number of channels;
[0080] Interleave dynamic convolutions (C2f-DynamicConv) among multiple downsampling convolutions in the backbone network to extract features of different scales, and finally enter the SPPF module to output the feature results, completing the feature extraction process of the backbone network. By interleaving multiple C2f-DynamicConv in the backbone network, rich features can be extracted more effectively, enabling it to better adapt to different inputs and improving the efficiency of the model. As the network deepens, the size of the feature map gradually decreases while the number of channels gradually increases, and finally enters the SPPF module in the YOLOv8 model. The SPPF module uses a multi-scale max pooling structure to enhance the receptive field and integrate multi-scale context information without changing the size of the feature map, improving the robustness to different target sizes.
[0081] It should be noted that in the feature fusion stage of the neck network, it includes:
[0082] Perform upsampling convolution on the feature results output by the backbone network, and perform a concatenation operation with the features of different scales output by the dynamic convolution (C2f-DynamicConv) in the backbone network to obtain the concatenated scale features;
[0083] Fuse the different scale features after concatenation according to the dynamic convolution (C2f-DynamicConv) to form a rich multi-scale semantic information flow. Structurally, the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) are adopted, enabling the model to more efficiently fuse the improved combination strategy of feature maps of different scales. First, it gradually upsamples from top to bottom and fuses with the shallow features of the corresponding layer, and then enhances the high-level semantic expression from bottom to top through downsampling. After each concatenation, a C2f_DynamicConv module is used to further refine the information to ensure that the fused feature map contains both fine spatial details and abstract semantic features. Through this two-way path enhancement mechanism, the model can balance the detection capabilities of small, medium, and large targets.
[0084] It should be noted that in S300, a partial convolution strategy is introduced in each detection head to reduce the computational complexity of the model detection head, specifically:
[0085] In the neck network feature fusion stage, multiple groups of fused feature maps are respectively input into the detection head for object detection. In this embodiment, there are three groups of feature maps. The detection head adopts a decoupled head structure, which separately processes bounding box regression and class prediction to improve the prediction accuracy, and introduces partial convolution to reduce the model calculation amount. Feature maps of each scale respectively correspond to small, medium, and large objects in the detected image. After passing through the detection head, detection boxes, confidence levels, and class information are output. Finally, all candidate bounding boxes pass through the detection head to remove redundancy and output the final detection results.
[0086] It should be noted that in S500, the spatial transformation network includes a localization network, a grid generator, and a sampler, where the localization network is a convolutional neural network;
[0087] Specifically, correcting the shape of the recognition dataset according to the spatial transformation network includes:
[0088] Using the localization network to receive the input image in the invoice text information recognition dataset for predicting the positions of a group of key control points in the input image, where the key control points represent the regions in the image that need to be deformed;
[0089] Using the grid generator to generate a deformation grid according to the positions of the key control points by using the thin plate spline interpolation algorithm, and defining the new positions of each pixel in the input image in the output corrected image according to the deformation grid;
[0090] Using the sampler to sample the input image according to the deformation grid to generate the output corrected image, making the text arrangement more horizontal and neat.
[0091] It should be noted that the spatial transformation network is a learnable module, which can automatically learn the spatial transformation parameters of the input image to correct the image. Traditional spatial transformation networks usually use affine transformation, which has limited effect on processing curved or irregular deformations. The spatial transformation network based on thin plate spline interpolation (TPS) adopts a powerful non-linear transformation method, which can better fit various complex deformations. Therefore, introducing the idea of thin plate spline interpolation into the spatial transformation network and using the thin plate spline interpolation transformation to replace the affine transformation in the spatial transformation network can enable the spatial transformation network to better handle various text deformations in taxi invoice images and more accurately locate and correct text regions.
[0092] It should be noted that the localization network is usually a convolutional neural network for predicting the set of fiducial points in the input image , where k is a constant. In this embodiment, k is 20, representing the input image Points in the upper normalized coordinate system, with its origin at the center of the image, so the abscissa and the ordinate both range from [-1, 1]. These fiducial points encode the geometric deformations (such as bending, tilting) of the input image and serve as the "anchor points" of the TPS transformation, determining how the input image is deformed to match a predefined set of base fiducial points, thereby establishing a pixel-level mapping relationship for precise correction. Since the TPS transformation is differentiable, the entire STN can be trained end-to-end, optimizing the localization network through backpropagation to make it predict better control points.
[0093] The grid generator is responsible for constructing the transformation mapping for image resampling. Its core function is to generate a dense pixel-level transformation grid based on the set of fiducial points predicted by the localization network through the TPS method. The grid generator first receives K fiducial points predicted by the localization network. These fiducial points capture the text shape features in the input image . At the same time, the grid generator predefines K base fiducial points C′, which are evenly distributed on the top and bottom edges of the corrected image to form a regular shape. Then, the parameter matrix T of the TPS transformation is calculated. This matrix defines the mapping relationship from the corrected image to the input image , and the calculation formula is as follows:
[0094]
[0095] In the formula, is a constant matrix of size determined only by the base fiducial points , which contains the geometric relationships between the base fiducial points. The representation of is as follows:
[0096]
[0097] In the formula, and are all-ones vectors used to handle translation, is the base fiducial point coordinate matrix used to handle scaling and rotation, is a matrix of size , where each element in the matrix is defined as:
[0098]
[0099] In the formula is the base fiducial point and the base reference point The Euclidean distance between them
[0100] The corrected image The pixel grid on it is represented as , where are the coordinates of the i-th pixel, and N is the total number of pixels
[0101] For the point on the corrected image , calculate its corresponding point on the input image , and the calculation formula is as follows
[0102]
[0103] In the formula is the abscissa of the i-th pixel point in the corrected image , is the ordinate of the i-th pixel point in the corrected image , represents the i-th pixel point in the corrected image and the thin plate spline kernel function value between the k-th base reference point ; By calculating for all pixel points on the corrected image as , a pixel grid can be obtained, and this pixel grid defines the mapping relationship between each pixel in the output corrected image and the corresponding position in the input image . However, these corresponding positions are usually not integer coordinates, so the bilinear interpolation method is used to determine the specific value of each pixel in the output corrected image .
[0104] The sampler is used to perform the interpolation process, sample the correct pixel value from the input image , and fill it into the output corrected image , which can be expressed as
[0105]
[0106] In the formula represents the corrected image, V represents the bilinear interpolation sampler represents the pixel grid generated by the grid generator represents the input image
[0107] It should be noted that in S500, enhancing the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR specifically includes:
[0108] Receiving the output corrected image and dividing it in the image patch embedding module to obtain image patches of a fixed size;
[0109] Performing convolutional mapping based on the image patches of the fixed size to obtain an initial feature sequence of the output corrected image;
[0110] Passing the initial feature sequence through a first mixing module and a second mixing module to generate a compressed feature sequence. Both the first mixing module and the second mixing module contain multiple local mixings. This module mainly combines the local window self-attention mechanism and the feed-forward network to implement local feature modeling and fusion operations, which are used to gradually compress the length of the feature sequence, improve the calculation efficiency and enhance the semantic abstraction ability;
[0111] Passing the compressed feature sequence through a third mixing module to generate a final feature sequence. The third mixing module includes one local mixing and five global mixings;
[0112] Passing the final feature sequence through a fully connected layer for character classification to output the recognition result.
[0113] It should be noted that in the SVTR, local mixing and global mixing are two important feature processing strategies, which can be weighed and adjusted according to specific application scenarios and input features. Global Mixing focuses on capturing the long-range dependencies between characters, thereby helping the model understand the context information. For curved or tilted deformed text, Global Mixing can also significantly improve the recognition performance by capturing global features. On the other hand, Local Mixing pays more attention to fine-grained local features, especially applicable to scenarios where the text area is dense and the overall layout is relatively fixed. In the invoice image, for a small text area with strong inter-character dependence and a compact structure, such as Chinese characters or fields with small letter spacing, such as invoice codes or amounts, Local Mixing can effectively capture the local relationships between characters and phrases and extract fine-grained features. In the case of a complex background or local noise, such as a complex bill background or slight blur, Local Mixing can effectively separate the local useful information and enhance the robustness of the model.
[0114] The composition of the Mixing Block in the original SVTR-S model is ((L, L, L), (L, L, L, L, L, G), (G, G, G, G, G, G)), where L represents Local Mixing and G represents Global Mixing. Considering the fields to be recognized, such as date, boarding and alighting times, unit price, mileage, amount, and fuel surcharge, etc., none of them are long text fields that rely on global semantics, but structured or semi-structured data, mainly relying on local features and format rules for recognition. For example, the date, boarding time, and alighting time usually adopt fixed formats, and the recognition depends on the combination of numbers and delimiters. The unit price, mileage, amount, and fuel surcharge are all numerical fields, usually with units, and the recognition mainly depends on the combination of numbers and units, none of which rely on global semantics. Although Global Mixing can also capture these local information to a certain extent, its main advantage lies in processing long texts and capturing long-distance dependencies. For these structured fields, the efficiency is not as high as Local Mixing. Therefore, in order to strengthen the model's ability to capture fine-grained information in these fields, the mixing block in SVTR is optimized by appropriately increasing the proportion of local mixing. After experimental verification, the composition of the new mixing block is finally adjusted to ((L, L, L), (L, L, L, L, L, L), (L, G, G, G, G, G)). This adjustment strategy significantly enhances the model's perception of local details while retaining a certain ability to capture global context information, making it more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, thus more effectively recognizing various structured fields on invoices. The present invention improves the Mixing Block by adjusting the weight ratio of Local Mixing and Global Mixing in SVTR to enhance the efficiency of taxi invoice image text recognition.
[0115] In this embodiment, the following specific implementation cases are provided to verify the text information recognition effect of the present invention, including the following two parts:
[0116] I. Verification of the present invention for improving the YOLOv8 model to enhance text detection effect
[0117] An existing taxi invoice detection dataset is used. This detection dataset contains 685 images, among which 479 images are selected as the training set and 206 images are selected as the test set. The image input size is set to [3, 640, 640]. To ensure the stability and repeatability of the experimental results, the same parameter configuration is used for all experiments, and the training is based on the deep learning model YOLOv8n.
[0118] To evaluate the performance impact of dynamic convolution and lightweight detection heads on the YOLOv8n model for taxi invoice text detection tasks, the present invention conducted a series of ablation experiments. The ablation experiment results for the improved model are shown in Table 1:
[0119] Table 1. Comparison of experimental results of the YOLOv8n model
[0120]
[0121] These experiments aimed to examine the effects of using these two optimization strategies individually and in combination. Using the YOLOv8n model as the baseline model, three variant models were constructed:
[0122] Model 1, which introduced dynamic convolution based on the baseline model;
[0123] Model 2, which performed lightweight transformation on the detection head of the baseline model, that is, used a lightweight detection head;
[0124] The present invention introduced dynamic convolution based on the baseline model and performed lightweight transformation on the detection head.
[0125] According to the experimental results:
[0126] For Model 1 with dynamic convolution introduced, there were improvements in mean average precision, precision, and recall. The mean average precision increased by 0.7%, precision increased by 1.0%, and recall increased by 0.3%, confirming that dynamic convolution can effectively improve detection accuracy. However, this improvement also led to an increase in the number of model parameters from 3.0M to 4.4M, while the computational cost decreased from 8.2G to 7.0G. This indicates that while dynamic convolution increases the parameters to some extent, it actually improves the computational efficiency when enhancing performance.
[0127] For Model 2 with a lightweight detection head, the number of model parameters and computational cost were significantly reduced. The number of parameters decreased from 3.0M to 2.4M, and the computational cost decreased from 8.2G to 5.6G, successfully achieving the goal of model lightweighting. However, at the same time, this strategy also led to a decrease in mean average precision, precision, and recall, which decreased by 0.3%, 0.2%, and 0.7% respectively, indicating that the lightweight operation sacrificed detection accuracy to some extent.
[0128] While the present invention maintained a relatively low number of parameters of 3.8M and computational cost of 4.4G, it still achieved a mean average precision of 93.6%, which is better than the baseline model, fully demonstrating that the present invention achieved a good balance between accuracy and efficiency.
[0129] II. Verification of the present invention for improving SVTR to enhance text recognition effect
[0130] Using the existing taxi invoice recognition dataset, which contains 3,959 text images. Among them, 3,520 text images are used as the training set for training, and 439 text images are used as the test set for testing. The input size of the text images is set to [3, 48, 256]. To ensure the stability and repeatability of the experimental results, the same hyperparameter configuration is used for all experiments. The training is based on the deep learning model SVTR-S (SVTR-Small), and a data augmentation strategy is used during the model training process. Subsequent data verification experiments are all carried out on the basis of data augmentation. A series of experiments are carried out after adding STN to the improved SVTR of the present invention and adjusting the proportion of the number of local mixings in SVTR. The text recognition accuracies before and after the experimental results are shown in Table 2:
[0131] Table 2. Comparison of experimental results of improving the proportion of the number of local mixings in the SVTR
[0132]
[0133] In the table, the original SVTR-S model consists of 8 Local Mixing and 7 Global Mixing;
[0134] Model 1 means that the Mixing Block consists of 8 Local Mixing and 7 Global Mixing;
[0135] Model 2 means that the Mixing Block consists of 9 Local Mixing and 6 Global Mixing;
[0136] Model 3 means that the Mixing Block consists of 10 Local Mixing and 5 Global Mixing;
[0137] Model 4 means that the Mixing Block consists of 11 Local Mixing and 4 Global Mixing;
[0138] Model 5 means that the Mixing Block consists of 12 Local Mixing and 3 Global Mixing;
[0139] Model 6 means that the Mixing Block consists of 13 Local Mixing and 2 Global Mixing;
[0140] Model 7 means that the Mixing Block consists of 14 Local Mixing and 1 Global Mixing.
[0141] As can be seen from the table, increasing the proportion of Local Mixing can improve the recognition accuracy of the model to a certain extent. Using Model 3 as the optimal embodiment of the present invention achieved the best result, with an accuracy rate of 82.69%, which is 2.51% higher than the original model. This shows that appropriately increasing the Local Mixing module can more effectively extract the local features of structured fields in scenarios such as invoices, verifying the effectiveness of the previous analysis. However, when continuing to increase the number of Local Mixing, such as from Model 4 to Model 7, the accuracy rate begins to fluctuate or even decline. This may be because when the number of Local Mixing is too large, the model will overfit to local details and ignore the feature combinations or structural information in a slightly larger range, thus reducing the recognition accuracy.
[0142] According to the invoice text information recognition method of the present invention, by dynamically generating the convolution kernel weights, the optimized YOLOv8 model can adaptively adjust the convolution operation according to the input, thereby enhancing the expression ability of the model and improving the text detection ability. In addition, by replacing the ordinary convolution in the YOLOv8 detection head with partial convolution (PConv), the amount of calculation and the number of parameters are significantly reduced, thus reducing the complexity of the model and improving the text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction. Utilizing the effectiveness of STN in geometric correction, it effectively alleviates geometric deformations such as text distortion and tilt in the input image, thereby improving the text recognition efficiency. By increasing the proportion of the local mixing in the SVTR to enhance the local feature perception ability and text information capture ability of the SVTR, it makes it more suitable for scenarios with relatively fixed formats but complex local details such as invoices, and thus more effectively recognizes various structured fields on the invoices.
[0143] Embodiment 2
[0144] This embodiment provides an invoice text information recognition system, in which the invoice text information is recognized by using the above recognition method. The system includes:
[0145] An acquisition module, configured to acquire an invoice image as a detection data set and label the key text in the invoice image to obtain training samples;
[0146] A first construction module, configured to construct a YOLOv8 model based on the dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and a plurality of detection heads. The dynamic convolution in ParameterNet is introduced into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model;
[0147] A training output module, which is used to introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model's detection head, improve the inference speed and resource utilization efficiency of the model, enhance the text detection ability, and train the YOLOv8 model according to the training samples to output the text regions of the training samples;
[0148] A second construction module, which is used to construct an invoice text information recognition dataset according to the text regions;
[0149] A third construction module, which is used to construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; the shape of the recognition dataset is corrected according to the spatial transformation network, and the local feature perception ability and text information capture ability of the SVTR are enhanced by increasing the proportion of the number of local mixing in the SVTR;
[0150] An output module, which is used to input the invoice text information recognition dataset into the text recognition model to output structured invoice text information.
[0151] An invoice text information recognition system in an embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0152] An invoice text information recognition system in an embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0153] An invoice text information recognition system provided by an embodiment of the present application can implement Figure 1 each process implemented by the method embodiment of a method for recognizing invoice text information. To avoid repetition, it will not be elaborated here.
[0154] The invoice text information recognition system according to the present invention dynamically generates convolution kernel weights, enabling the optimized YOLOv8 model to adaptively adjust convolution operations according to the input, thereby enhancing the model's expressive ability and improving text detection ability. In addition, by replacing ordinary convolutions in the detection head of the YOLOv8 model with partial convolutions (PConv), the amount of computation and the number of parameters are significantly reduced, thus reducing the complexity of the model and improving text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction. By utilizing the effectiveness of STN in geometric correction, geometric deformations such as text distortion and inclination existing in the input image are effectively alleviated, thereby improving text recognition efficiency. By increasing the proportion of the local mixture in SVTR, the local feature perception ability and text information capture ability of SVTR are enhanced, making it more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, so as to more effectively identify various structured fields on invoices.
[0155] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the invoice text information recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0156] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the invoice text information recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0157] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0158] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention.
[0159] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.
[0160] Obviously, the described embodiments are only a part of the embodiments of this application, rather than all embodiments. The mention of "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0161] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for identifying invoice text information, characterized in that, Including: Obtain an invoice image as a detection dataset, and annotate the key texts in the invoice image to obtain training samples; Construct a YOLOv8 model based on dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. Introduce the dynamic convolution in ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; Introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and train the YOLOv8 model according to the training samples to output the text regions of the training samples; Construct an invoice text information recognition dataset according to the text regions; Construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; Correct the shape of the recognition dataset according to the spatial transformation network, and enhance the local feature perception ability and text information capture ability of the scene text recognition network based on a single vision model by increasing the proportion of the number of local mixing in the scene text recognition network based on a single vision model; Input the invoice text information recognition dataset into the text recognition model to output structured invoice text information.
2. The invoice text information recognition method according to claim 1, characterized in that The spatial transformation network includes a localization network, a grid generator, and a sampler, where the localization network is a convolutional neural network; Correcting the shape of the recognition dataset according to the spatial transformation network specifically includes: Use the localization network to receive the input image in the invoice text information recognition dataset, and predict the positions of a set of key control points in the input image, where the key control points represent the regions in the image that need to be deformed; Use the grid generator to generate a deformation grid according to the positions of the key control points, and define the new positions of each pixel in the input image in the output corrected image according to the deformation grid; Use the sampler to sample the input image according to the deformation grid to generate the output corrected image.
3. The invoice text information recognition method according to claim 1, characterized in that Enhancing the local feature perception ability and text information capture ability of the scene text recognition network based on a single vision model by increasing the proportion of the number of local mixing in the scene text recognition network based on a single vision model specifically includes: Receive the output corrected image to the image patch embedding module for division to obtain image patches of a fixed size; Perform convolutional mapping according to the image patches of the fixed size to obtain the initial feature sequence of the output corrected image; Pass the initial feature sequence through a first mixing module and a second mixing module to generate a compressed feature sequence. Both the first mixing module and the second mixing module contain multiple local mixings; Pass the compressed feature sequence through a third mixing module to generate the final feature sequence. The third mixing module includes one local mixing and five global mixings; The final feature sequence is passed through a fully connected layer for character classification to output the recognition result.
4. The invoice text information recognition method according to claim 1, wherein The dynamic convolution network based on ParameterNet is a network that generates convolution kernels and weights by combining multiple dynamic experts, through multiple predefined convolution kernels and the dynamic coefficients corresponding to the convolution kernels which are generated by weighted combination; Output Feature Map of YOLOv8 Model Based on Dynamic Convolution Network The calculation formula is as follows: Among them, represents the input feature map, and its dimensions are expressed as , represents the number of input channels, H represents the input height, and W represents the input width; represents the output feature map, which is the output result of the dynamic convolution operation, and its dimensions are expressed as , represents the number of output channels, represents the output height, represents the output width, represents the i-th dynamic expert, and M represents the number of dynamic experts; Dynamic coefficient is generated by global average pooling and a multi-layer perceptron based on input features, and the generation formula is as follows: Pool represents global average pooling, and MLP represents multi-layer perceptron.
5. The invoice text information recognition method according to claim 1, characterized in that, During the annotation process, rectangular bounding boxes are used to select the key texts of each invoice image to obtain the corresponding target label file, thereby obtaining the training samples.
6. The invoice text information recognition method according to claim 1, wherein In the feature extraction stage of the backbone network, it includes: the training samples pass through multiple downsampling convolution operations to extract multi-scale features, dynamic convolutions are interspersed among multiple downsampling convolutions in the backbone network to extract different-scale features, and finally enter the SPPF module in the YOLOv8 model to output the feature results, completing the feature extraction process of the backbone network.
7. The invoice text information recognition method according to claim 1, characterized in that In the feature fusion stage of the neck network, it includes: Performing upsampling convolution on the feature results output by the backbone network, and performing a splicing operation with different-scale features output by the dynamic convolution in the backbone network to obtain the spliced scale features; Fusing the spliced different-scale features according to the dynamic convolution to form a multi-scale semantic information flow.
8. An invoice text information recognition system, characterized in that, Using the recognition method described in any one of claims 1 to 7 to recognize the invoice text information, the system includes: An acquisition module, configured to acquire an invoice image as a detection data set, and annotate the key texts in the invoice image to obtain training samples; A first construction module, configured to construct a YOLOv8 model based on the dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. The dynamic convolution in ParameterNet is introduced into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; A training output module, configured to introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and train the YOLOv8 model according to the training samples to output the text regions of the training samples; A second construction module, configured to construct an invoice text information recognition data set according to the text regions; A third construction module, configured to construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; correcting the shape of the recognition data set according to the spatial transformation network, and enhancing the local feature perception ability and text information capture ability of the scene text recognition network based on a single vision model by increasing the proportion of the number of local mixing in the scene text recognition network based on a single vision model; An output module, configured to input the invoice text information recognition data set into the text recognition model to output the structured invoice text information.
9. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the invoice text information recognition method described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The computer storage medium stores instructions, which, when executed on a computer, cause the computer to execute the invoice text information recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Bill information identification method, device and equipment based on OCR (Optical Character Recognition) and storage medium
CN118397642A
Value-added tax invoice image content segmentation method and system based on YOLOv8-Seg
CN119314194A