Invoice text information identification method and system

By introducing dynamic convolution and partial convolution strategies in the YOLOv8 model, combined with the TPS spatial transformation network and the local hybrid enhanced text recognition network, the problems of insufficient robustness of image preprocessing, limited text detection accuracy and accumulated recognition errors in taxi invoice text recognition are solved, and more efficient and accurate text information recognition is achieved.

CN120088810AActive Publication Date: 2025-06-03NANCHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510577886.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When handling complex scenarios such as taxi invoices, the prior art has problems such as insufficient robustness of image preprocessing, limited text detection accuracy, and accumulated text recognition errors.

Method used

The YOLOv8 model based on ParameterNet is adopted, combined with partial convolution strategies, and the feature extraction capability and the calculation efficiency of the detection head are enhanced. At the same time, a spatial transformation network based on TPS and a local hybrid enhanced scene text recognition network are introduced to perform text image correction and feature perception.

Benefits of technology

It improves the accuracy and robustness of text detection, reduces the complexity and calculation cost of the model, and enhances the ability to identify invoice text information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088810A_ABST
    Figure CN120088810A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of invoice identification, in particular to an invoice text information identification method and system. Comprising the following steps: constructing a YOLOv8 model based on dynamic convolution of Parameter Net, and introducing the dynamic convolution in the Parameter Net into C2f modules in a backbone network and a neck network so as to enhance the feature extraction capability of the YOLOv8 model; introducing a partial convolution strategy to each detection head to reduce the calculation complexity of the model detection head, and training the YOLOv8 model according to the training sample to output a text region; and constructing an STN including a TPS and an SVTR based on a single vision model, correcting the shape of the recognition data set according to the STN, enhancing the local feature perception capability and the text information capture capability by increasing the number proportion of the local mixture in the SVTR, and finally outputting structured invoice text information. According to the method, the expression ability of the model is enhanced, the text detection ability and recognition efficiency are improved, and the local feature perception ability and the text information capture ability of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of invoice recognition, and in particular to a method and system for recognizing invoice text information. Background Art

[0002] With the development of information technology, financial reimbursement has gradually stepped onto the road of informatization. The traditional manual entry of bill information will be gradually replaced by intelligent recognition. The electronicization of bills and the text detection and recognition of electronic bills directly determine the feasibility of intelligent financial reimbursement. As one of the many bills, taxi invoices have their own uniqueness and complexity in image text detection and recognition due to the particularity of their layout format.

[0003] In real life, the text layout of original taxi invoices is diverse. Some text lines have different spacing, some text blocks are irregularly distributed, some text is dense, resulting in a large amount of text information concentrated in a limited area, and some invoice image quality is uncertain, such as image blur, uneven lighting, and stains. These factors affect the accuracy and robustness of text detection. Traditional OCR technology relies on manual feature extraction and template matching, so it is difficult to deal with problems such as dense text, blurred fonts, tilted deformation, and complex background interference in taxi invoice images.

[0004] Some existing end-to-end models based on deep learning (such as YOLOv8, DBNet, CRNN, etc.) have gradually become the mainstream solutions for invoice text recognition. However, they still have the following key defects when dealing with complex scenarios such as taxi invoices: 1. Insufficient robustness of image preprocessing. Taxi invoices are often seriously disturbed by factors such as shooting environment and storage conditions. The uneven lighting problem makes it difficult to accurately extract the text area, and there is interference from folding and staining.

[0005] 2. Limited text detection accuracy. In the densely packed area of ​​taxi invoices, the spacing between text lines is extremely small, which makes it easy for anchor-based detection methods to make overlapping frame merging errors. The small size of characters in key information such as invoice codes makes lightweight models (such as YOLOv5s) have insufficient receptive fields, resulting in significant positioning offsets. Some tilted and curved texts make mainstream horizontal rectangular box detection unsuitable and must rely on additional angle regression modules (such as RRPN), which increases computational costs.

[0006] 3. Accumulation of text recognition errors. The text recognition stage is limited by font diversity, low resolution, and semantic relevance. The complex fonts of taxi invoices, such as characters with similar shapes (such as "3" and "8", etc.), can easily lead to misjudgment. Low-quality images such as blurred and faded texts cause convolutional feature extraction to fail, and recognition is prone to errors. Summary of the invention

[0007] The present invention aims to at least improve one of the technical problems existing in the prior art. To this end, the present invention proposes an invoice text information recognition method and system.

[0008] The technical solution of the present invention is as follows: An invoice text information recognition method, which includes: Obtain an invoice image as a detection data set, and annotate key texts in the invoice image to obtain training samples; Construct a YOLOv8 model based on the dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. Introduce the dynamic convolution in ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; Introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and train the YOLOv8 model according to the training samples to output the text area of the training samples; Construct an invoice text information recognition data set according to the text area; Construct a text recognition model. The text recognition model includes a spatial transformation network (STN) based on thin plate spline interpolation (TPS) and a scene text recognition network (SVTR) based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; Correct the shape of the recognition data set according to the spatial transformation network, and enhance the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR; Input the invoice text information recognition data set into the text recognition model to output structured invoice text information.

[0009] In a possible technical solution, further, the spatial transformation network includes a positioning network, a grid generator, and a sampler, where the positioning network is a convolutional neural network; Specifically, correcting the shape of the recognition data set according to the spatial transformation network includes: Use the positioning network to receive the input image in the invoice text information recognition data set, and predict the positions of a set of key control points in the input image, where the key control points represent the areas in the image that need to be deformed; Use the grid generator to generate a deformation grid according to the positions of the key control points using the thin plate spline interpolation algorithm, and define the new positions of each pixel in the input image in the output corrected image according to the deformation grid; The sampler samples the input image according to the deformed grid to generate an output corrected image, making the text arrangement more horizontal and neat.

[0010] In a possible technical solution, further, enhancing the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixing in the SVTR specifically includes: Receiving the output corrected image to the image block embedding module for division to obtain image blocks of a fixed size; Performing convolutional mapping according to the image blocks of the fixed size to obtain an initial feature sequence of the output corrected image; Passing the initial feature sequence through a first mixing module and a second mixing module to generate a compressed feature sequence. Both the first mixing module and the second mixing module contain multiple local mixings. The first mixing module and the second mixing module mainly combine the local window self-attention mechanism and the feed-forward network to implement local feature modeling and fusion operations, for gradually compressing the length of the feature sequence, improving the calculation efficiency and enhancing the semantic abstraction ability; Passing the compressed feature sequence through a third mixing module to generate a final feature sequence. The third mixing module includes one local mixing and five global mixings; Passing the final feature sequence through a fully connected layer for character classification to output the recognition result.

[0011] In a possible technical solution, further, the dynamic convolutional network based on ParameterNet is generated by combining the convolutional kernels and weight generation networks of multiple Dynamic Experts (DEs), through multiple predefined convolutional kernels and the dynamic coefficients corresponding to the convolutional kernels obtained through weighted combination; The output feature map of the YOLOv8 model based on the dynamic convolutional network The calculation formula is: Among them, represents the input feature map, and its dimension is expressed as , represents the number of input channels, H represents the input height, and W represents the input width; represents the output feature map, which is the output result of the dynamic convolutional operation, and its dimension is expressed as , represents the number of output channels, represents the output height, , represents the i-th dynamic expert, and M represents the number of dynamic experts; Dynamic coefficient is generated by global average pooling and a multi-layer perceptron based on input features, and the generation formula is: Pool represents global average pooling, and MLP represents a multi-layer perceptron. Dynamic convolution can dynamically generate convolution kernel parameters according to input features, so as to better adapt to text regions of different shapes and sizes, and enhance the robustness of the model to deformation and noise. At the same time, the parameterization mechanism provided by ParameterNet can effectively control the computational complexity of dynamic convolution, so that while maintaining high performance, it will not significantly increase the computational burden of the model, and thus is more suitable for resource-constrained practical application scenarios.

[0012] In a possible technical solution, further, during the annotation process, rectangular bounding boxes are used to frame the key text of each invoice image to obtain the corresponding target label file, thereby obtaining training samples; wherein, when framing, the framing range can completely cover the key text, while minimizing the interference of irrelevant backgrounds, so as to improve the accuracy and robustness of subsequent model training.

[0013] In a possible technical solution, further, in the feature extraction stage of the backbone network, it includes: The training samples are subjected to multiple downsampling convolution operations to extract multi-scale features, where the convolution kernel size of the downsampling convolution is 3 and the stride is 2, which is used to reduce the size of the feature map and increase the number of channels; Dynamic convolutions (C2f-DynamicConv) are interspersed among multiple downsampling convolutions in the backbone network to extract different-scale features, and finally enter the SPPF module in the YOLOv8 model to output feature results, completing the backbone network feature extraction process. By interspersing multiple C2f-DynamicConv in the backbone network, richer features can be extracted more effectively, enabling it to better adapt to different inputs and improving the efficiency of the model. As the network deepens, the size of the feature map gradually decreases while the number of channels gradually increases, and finally enters the SPPF module. The SPPF module uses a multi-scale max-pooling structure to enhance the receptive field and integrate multi-scale context information without changing the size of the feature map, improving the robustness to different target sizes.

[0014] In a possible technical solution, further, in the neck network feature fusion stage, it includes: Upsampling convolution is performed on the feature results output by the backbone network, and splicing operations are performed with different-scale features output by the dynamic convolution (C2f-DynamicConv) in the backbone network to obtain the spliced scale features; Fuse the spliced features of different scales according to the dynamic convolution (C2f-DynamicConv) to form a rich multi-scale semantic information flow.

[0015] According to the invoice text information recognition method of the present invention, by dynamically generating the weights of the convolution kernel, the optimized YOLOv8 model can adaptively adjust the convolution operation according to the input, thereby enhancing the expression ability of the model and improving the text detection ability. In addition, by replacing the ordinary convolution in the YOLOv8 detection head with partial convolution (PConv), the amount of calculation and the number of parameters are significantly reduced, thereby reducing the complexity of the model and improving the text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction, and utilizes the effectiveness of STN in geometric correction to effectively alleviate geometric deformations such as text distortion and tilt in the input image, thereby improving the text recognition efficiency. By increasing the proportion of the local mixing in the SVTR to enhance the local feature perception ability and text information capture ability of the SVTR, it is more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, so as to more effectively identify various structured fields on the invoices.

[0016] An invoice text information recognition system, wherein the above recognition method is used to recognize the invoice text information, and the system includes: An acquisition module, configured to acquire an invoice image as a detection data set, and label the key text in the invoice image to obtain training samples; A first construction module, configured to construct a YOLOv8 model based on the dynamic convolution of ParameterNet, the YOLOv8 model includes a backbone network, a neck network, and a plurality of detection heads, and introduce the dynamic convolution in the ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; A training output module, configured to introduce a partial convolution strategy into each detection head to reduce the computational complexity of the model detection head, and train the YOLOv8 model according to the training samples to output the text area of the training samples; A second construction module, configured to construct an invoice text information recognition data set according to the text area; A third construction module is used to construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on the single vision model includes local mixing and global mixing. According to the spatial transformation network, the shape of the recognition data set is corrected, and the local feature perception ability and text information capture ability of the SVTR are enhanced by increasing the proportion of the local mixing in the SVTR. An output module is used to input the invoice text information recognition data set into the text recognition model to output structured invoice text information.

[0017] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the invoice text information recognition method as described above is implemented.

[0018] A computer storage medium stores instructions, and when the instructions are executed on a computer, the computer is made to execute the invoice text information recognition method as described above.

[0019] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of the invoice text information recognition method according to an embodiment of the present invention; Figure 2 It is an architecture diagram of the text recognition model of the invoice text information recognition method according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the invoice text information recognition system according to an embodiment of the present invention. Detailed Embodiments

[0022] The embodiments of the present invention will be described in detail below. The embodiments described with reference to the drawings are exemplary. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0023] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected to" another element, it can be directly connected to the other element or there may be an intermediate element at the same time.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0025] The terms "first", "second", "third", etc. in the description and claims of this application and the accompanying drawings are used to distinguish different objects and are not used to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, steps or units not listed are also included, or optionally, other steps or units inherent to these processes, methods, products or devices are also included.

[0026] Only parts relevant to this application are shown in the drawings, not all of the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0027] The terms "component", "module", "system", "unit", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or distributed between two or more computers. In addition, these units can be executed from various computer-readable media storing various data structures. A unit can communicate, for example, through signals with other systems through local and / or remote processes according to signals having one or more data packets (such as data from a second unit interacting with a local system, a distributed system, and / or a network. For example, the Internet interacting with other systems through signals).

[0028] Embodiment 1 This embodiment provides a method for identifying invoice text information, which includes: S100, obtaining an invoice image as a detection data set, and annotating key texts in the invoice image to obtain training samples; S200, constructing a YOLOv8 model based on dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. Introduce the dynamic convolution in ParameterNet into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; S300, introducing a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, improve the inference speed and resource utilization efficiency of the model, and improve the text detection ability. Train the YOLOv8 model according to the training samples to output the text regions of the training samples; S400, constructing an invoice text information recognition data set according to the text regions; S500, constructing a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; Correct the shape of the recognition data set according to the spatial transformation network, and enhance the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the number of local mixing in the SVTR; S600, inputting the invoice text information recognition data set into the text recognition model to output structured invoice text information. It should be noted that in S100, during the annotation process, a rectangular bounding box is used to select the key texts of each invoice image to obtain the corresponding target label file, thereby obtaining training samples; Among them, when selecting, ensure that the selection range can completely cover the key texts, and at the same time minimize the interference of irrelevant backgrounds to improve the accuracy and robustness of subsequent model training. After annotation, each box corresponds to a target label, and using the LabelImg annotation tool will automatically generate the corresponding label file. The label file stores the annotation information in text format, with each line corresponding to a target label. The first number represents the target category, and the following four numbers represent the position information of the target bounding box, including the center point coordinates (x, y) of the bounding box and the width and height (w, h). These values are all normalized between the width and height of the image, and the value range is [0,1], which is convenient for subsequent deep learning models to read and use.

[0029] It should be noted that in S200, the dynamic convolution network based on ParameterNet is a network that generates convolution kernels and weights by combining multiple Dynamic Experts (DEs). It can dynamically adjust the convolution kernel weights to adapt to the characteristics of the input data, effectively capture the features of small text regions on invoices, and adapt to various font and layout forms. In addition, it can effectively suppress the interference of complex backgrounds such as invoice stamps, lines, or patterns, focus on the text region, and thus improve the robustness and accuracy of detection. By enhancing the model's adaptability to geometric changes and detailed features, dynamic convolution can significantly improve the accuracy and generalization performance of taxi invoice text detection.

[0030] In the improved YOLOv8 model of this embodiment, for each input feature map X, the model no longer uses a single fixed convolution kernel, but generates it through a weighted combination of multiple predefined convolution kernels and the dynamic coefficients corresponding to the convolution kernels ; The output feature map of the YOLOv8 model based on the dynamic convolution network is calculated by the formula: where, represents the input feature map, and its dimension is expressed as , represents the number of input channels, H represents the input height, and W represents the input width; represents the output feature map, which is the output result of the dynamic convolution operation, and its dimension is expressed as , represents the number of output channels, represents the output height, , represents the i-th dynamic expert, and M represents the number of dynamic experts; The dynamic coefficient is generated based on the input features through global average pooling and a multi-layer perceptron. The formula for generating is: Pool represents global average pooling, and MLP represents a multi-layer perceptron. Dynamic convolution can dynamically generate convolution kernel parameters according to the input features, so as to better adapt to text regions of different shapes and sizes, and enhance the model's robustness to deformation and noise. At the same time, the parameterization mechanism provided by ParameterNet can effectively control the computational complexity of dynamic convolution, so that while maintaining high performance, it will not significantly increase the computational burden of the model, and thus is more suitable for practical application scenarios with limited resources.

[0031] It should be noted that in this embodiment, in the backbone network feature extraction stage, it includes: The training samples pass through multiple downsampling convolution operations to extract multi-scale features. The convolution kernel size of the downsampling convolution is 3 and the stride is 2, which is used to reduce the size of the feature map and increase the number of channels; Dynamic convolutions (C2f-DynamicConv) are interspersed among multiple downsampling convolutions in the backbone network to extract features of different scales, and finally enter the SPPF module to output the feature results, completing the backbone network feature extraction process. By interspersing multiple C2f-DynamicConv in the backbone network, richer features can be extracted more effectively, enabling it to better adapt to different inputs and improving the efficiency of the model. As the network deepens, the size of the feature map gradually decreases while the number of channels gradually increases, and finally enters the SPPF module in the YOLOv8 model. The SPPF module uses a multi-scale max pooling structure to enhance the receptive field and integrate multi-scale context information without changing the size of the feature map, improving the robustness to different target sizes.

[0032] It should be noted that in the neck network feature fusion stage, it includes: Perform upsampling convolution on the feature results output by the backbone network, and perform a splicing operation with the features of different scales output by the dynamic convolution (C2f-DynamicConv) in the backbone network to obtain the spliced scale features; Fuse the spliced features of different scales according to the dynamic convolution (C2f-DynamicConv) to form a rich multi-scale semantic information flow. Structurally, the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) are adopted, enabling the model to more efficiently fuse the improved combination strategy of feature maps of different scales. First, gradually upsample from top to bottom and fuse with the shallow features of the corresponding layer, and then enhance the high-level semantic expression from bottom to top through downsampling. After each splicing, a C2f_DynamicConv module is used to further refine the information to ensure that the fused feature map contains both fine spatial details and abstract semantic features. Through this two-way path enhancement mechanism, the model can balance the detection capabilities for small, medium, and large targets.

[0033] It should be noted that in S300, a partial convolution strategy is introduced in each detection head to reduce the computational complexity of the model detection head, specifically: In the neck network feature fusion stage, multiple groups of fused feature maps are respectively input into the detection head for object detection. In this embodiment, there are three groups of feature maps. The detection head adopts a decoupled head structure, which separately processes bounding box regression and class prediction to improve the prediction accuracy, and introduces partial convolution to reduce the model's computational complexity. Feature maps of each scale respectively correspond to small, medium, and large objects in the detected image. After passing through the detection head, detection boxes, confidence levels, and class information are output. Finally, all candidate bounding boxes pass through the detection head to remove redundancy and output the final detection results.

[0034] It should be noted that in S500, the spatial transformation network includes a localization network, a grid generator, and a sampler, where the localization network is a convolutional neural network; Specifically, correcting the shape of the recognition dataset according to the spatial transformation network includes: Using the localization network to receive the input image in the invoice text information recognition dataset for predicting the positions of a set of key control points in the input image, where the key control points represent the areas in the image that need to be deformed; Using the grid generator to generate a deformation grid according to the positions of the key control points using the thin plate spline interpolation algorithm, and defining the new positions of each pixel in the input image in the output corrected image according to the deformation grid; Using the sampler to sample the input image according to the deformation grid to generate the output corrected image, making the text arrangement more horizontal and neat.

[0035] It should be noted that the spatial transformation network is a learnable module that can automatically learn the spatial transformation parameters of the input image to correct the image. Traditional spatial transformation networks usually use affine transformation, which has limited effect on processing curved or irregular deformations. The spatial transformation network based on thin plate spline interpolation (TPS) adopts a powerful non-linear transformation method that can better fit various complex deformations. Therefore, introducing the idea of thin plate spline interpolation into the spatial transformation network and using the thin plate spline interpolation transformation to replace the affine transformation in the spatial transformation network can enable the spatial transformation network to better handle various text deformations in taxi invoice images and more accurately locate and correct the text area.

[0036] It should be noted that the localization network is usually a convolutional neural network for predicting the set of fiducial points in the input image , where k is a constant. In this embodiment, k is 20, representing the points in the normalized coordinate system on the input image whose origin is located at the center of the image. Therefore, the abscissa and the ordinate The value range of is between [-1,1]. These reference points encode the input image The geometric deformation (such as bending, tilting) of the input image is determined as the "anchor point" of the TPS transformation. How to deform to match the predefined set of basic reference points, so as to establish a pixel-level mapping relationship and achieve accurate correction. Since the TPS transformation is differentiable, the entire STN can be trained end-to-end, and the positioning network can be optimized through back-propagation to predict better control points.

[0037] The grid generator is responsible for constructing the transformation map for image resampling. Its core function is to generate a dense pixel-level transformation grid based on the set of reference points predicted by the localization network through the TPS method. The grid generator first receives the K reference points predicted by the localization network. , these fiducials capture the input image At the same time, the grid generator predefines K basic reference points C′, which are evenly distributed in the rectified image. The top and bottom edges of the image form a regular shape. Then calculate the parameter matrix T of the TPS transformation, which defines the To the input image The mapping relationship is calculated as follows: In the formula, It is a base point Decided A constant matrix of size containing the geometric relationships between the base reference points. The representation is as follows: In the formula, and is a vector of all 1s, used to handle translation, is the base point coordinate matrix, used to handle scaling and rotation, is a The matrix of Each element in the matrix Defined as: In the formula Is the basic reference point and base reference points The Euclidean distance between .

[0038] Rectify the image The pixel grid on ,in is the coordinate of the i-th pixel, and N is the total number of pixels.

[0039] For the corrected image at the point , calculate its corresponding point on the input image , and the calculation formula is as follows: In the formula, is the abscissa of the i-th pixel point in the corrected image , is the ordinate of the i-th pixel point in the corrected image , represents the thin plate spline kernel function value between the i-th pixel point in the corrected image and the k-th basic reference point ; by calculating for all pixel points on the corrected image , a pixel grid can be obtained, and this pixel grid defines the mapping relationship between each pixel in the output corrected image and the corresponding position in the input image . However, these corresponding positions are usually not integer coordinates, so the bilinear interpolation method is used to determine the specific value of each pixel in the output corrected image .

[0040] The sampler is used to perform the interpolation process, sample the correct pixel values from the input image , and fill them into the output corrected image , which can be expressed as: In the formula, represents the corrected image, V represents the bilinear interpolation sampler, represents the pixel grid generated by the grid generator, represents the input image.

[0041] It should be noted that in S500, enhancing the local feature perception ability and text information capture ability of the SVTR by increasing the proportion of the local mixture in the SVTR specifically includes: Receiving the output corrected image and dividing it by the image block embedding module to obtain image blocks of a fixed size; Performing convolutional mapping based on the image blocks of the fixed size to obtain the initial feature sequence of the output corrected image; The initial feature sequence is passed through a first mixing module and a second mixing module to generate a compressed feature sequence. Both the first mixing module and the second mixing module contain multiple local mixings. This module mainly combines a local window self-attention mechanism with a feed-forward network to achieve local feature modeling and fusion operations, which are used to gradually compress the length of the feature sequence, improve computational efficiency, and enhance semantic abstraction ability. The compressed feature sequence is passed through a third mixing module to generate a final feature sequence. The third mixing module includes one local mixing and five global mixings. The final feature sequence is passed through a fully connected layer for character classification to output the recognition result.

[0042] It should be noted that in SVTR, local mixing and global mixing are two important feature processing strategies, which can be weighed and adjusted according to specific application scenarios and input features. Global Mixing focuses on capturing long-range dependencies between characters, thereby helping the model understand context information. For curved or skewed deformed text, Global Mixing can also significantly improve the recognition performance by capturing global features. On the other hand, Local Mixing pays more attention to fine-grained local features, especially suitable for scenarios where the text area is dense and the overall layout is relatively fixed. In invoice images, for small text areas with strong inter-character dependencies and compact structures, such as Chinese characters or fields with small letter intervals, such as invoice codes or amounts, Local Mixing can effectively capture the local relationships between characters and phrases and extract fine-grained features. In the case of complex backgrounds or local noise, such as complex bill backgrounds or slight blurring, Local Mixing can effectively separate local useful information and enhance the robustness of the model. The composition of the mixing block in the original SVTR-S model is ((L, L, L), (L, L, L, L, L, G), (G, G, G, G, G, G)), where L represents Local Mixing and G represents Global Mixing. Considering the fields to be recognized, such as date, boarding and alighting times, unit price, mileage, amount, and fuel surcharge, etc., none of them are long text fields that rely on global semantics, but structured or semi-structured data, and mainly rely on local features and format rules for recognition. For example, the date, boarding time, and alighting time usually adopt fixed formats, and the recognition depends on the combination of numbers and delimiters. The unit price, mileage, amount, and fuel surcharge are all numerical fields, usually with units, and the recognition mainly depends on the combination of numbers and units, none of which depend on global semantics. Although Global Mixing can also capture these local information to a certain extent, its main advantage lies in processing long texts and capturing long-distance dependencies. For these structured fields, the efficiency is not as high as Local Mixing. Therefore, in order to strengthen the model's ability to capture fine-grained information in these fields, the mixing block in SVTR was optimized, and the proportion of local mixing was appropriately increased. After experimental verification, the composition of the new mixing block was finally adjusted to ((L, L, L), (L, L, L, L, L, L), (L, G, G, G, G, G)). This adjustment strategy significantly enhances the model's perception of local details while retaining a certain ability to capture global context information, making it more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, thereby more effectively recognizing various structured fields on invoices. The present invention improves the Mixing Block by adjusting the weight ratio of Local Mixing and Global Mixing in SVTR to enhance the text recognition efficiency of taxi invoice images.

[0043] In this embodiment, the following specific implementation cases are provided to verify the text information recognition effect of the present invention, including the following two parts: I. Verification of the present invention for improving the YOLOv8 model to enhance text detection effect An existing taxi invoice detection dataset is used. This detection dataset contains 685 images, among which 479 images are selected as the training set and 206 images are selected as the test set. The image input size is set to [3, 640, 640]. To ensure the stability and repeatability of the experimental results, the same parameter configuration is used for all experiments, and the training is based on the deep learning model YOLOv8n.

[0044] To evaluate the performance impact of dynamic convolution and lightweight detection heads on the YOLOv8n model for taxi invoice text detection tasks, the present invention conducted a series of ablation experiments. The ablation experiment results for improving the model are shown in Table 1 as follows: Table 1. Comparison of experimental results of the YOLOv8n model These experiments aimed to examine the effects of using these two optimization strategies individually and in combination. Using the YOLOv8n model as the baseline model, three variant models were constructed: Model 1, which introduced dynamic convolution based on the baseline model; Model 2, which performed lightweight transformation on the detection head of the baseline model, that is, used a lightweight detection head; The present invention introduced dynamic convolution based on the baseline model and performed lightweight transformation on the detection head.

[0045] According to the experimental results: For Model 1 with dynamic convolution introduced, there were improvements in the mean average precision, precision, and recall. The mean average precision increased by 0.7%, the precision increased by 1.0%, and the recall increased by 0.3%, confirming that dynamic convolution can effectively improve detection accuracy. However, this improvement also led to an increase in the model's parameter count from 3.0M to 4.4M, while the computational cost decreased from 8.2G to 7.0G. This indicates that while dynamic convolution increases the parameters to some extent, it actually improves the computational efficiency while enhancing the performance.

[0046] Model 2 with a lightweight detection head significantly reduced the model's parameter count and computational cost. Its parameter count decreased from 3.0M to 2.4M, and the computational cost decreased from 8.2G to 5.6G, successfully achieving the goal of model lightweighting. However, at the same time, this strategy also led to a decrease in the mean average precision, precision, and recall, which decreased by 0.3%, 0.2%, and 0.7% respectively, indicating that the lightweight operation sacrificed detection accuracy to some extent.

[0047] While the present invention maintained a relatively low parameter count of 3.8M and computational cost of 4.4G, it still achieved a mean average precision of 93.6%, which is better than the baseline model, fully demonstrating that the present invention achieved a good balance between accuracy and efficiency.

[0048] II. Verification of the present invention for improving SVTR to enhance text recognition effect Using the existing taxi invoice recognition dataset, which contains 3,959 text images. Among them, 3,520 text images are used as the training set for training, and 439 text images are used as the test set for testing. The input size of the text images is set to [3, 48, 256]. To ensure the stability and repeatability of the experimental results, the same hyperparameter configuration is used for all experiments. The training is based on the deep learning model SVTR-S (SVTR-Small), and a data augmentation strategy is used during the model training process. Subsequent data verification experiments are all carried out on the basis of data augmentation. A series of experiments are carried out after adding STN to the improved SVTR of the present invention and adjusting the proportion of the number of local mixings in SVTR. The text recognition accuracies before and after the experimental results are shown in Table 2 as follows: Table 2. Comparison of experimental results of adjusting the proportion of the number of local mixings in the SVTR In the table, the original SVTR-S model consists of 8 Local Mixings and 7 Global Mixings; Model 1 means that the Mixing Block consists of 8 Local Mixings and 7 Global Mixings; Model 2 means that the Mixing Block consists of 9 Local Mixings and 6 Global Mixings; Model 3 means that the Mixing Block consists of 10 Local Mixings and 5 Global Mixings; Model 4 means that the Mixing Block consists of 11 Local Mixings and 4 Global Mixings; Model 5 means that the Mixing Block consists of 12 Local Mixings and 3 Global Mixings; Model 6 means that the Mixing Block consists of 13 Local Mixings and 2 Global Mixings; Model 7 means that the Mixing Block consists of 14 Local Mixings and 1 Global Mixing.

[0049] As can be seen from the table, increasing the proportion of Local Mixing can improve the recognition accuracy of the model to a certain extent. Using Model 3 as the optimal embodiment of the present invention achieved the best results, with an accuracy rate reaching 82.69%, an increase of 2.51% compared to the original model, indicating that appropriately increasing the Local Mixing module can more effectively extract the local features of structured fields in scenarios such as invoices, verifying the effectiveness of the previous analysis. However, when continuing to increase the number of Local Mixing, such as from Model 4 to Model 7, the accuracy rate began to fluctuate or even decline. This may be because when the number of Local Mixing is too large, the model will overfit to local details and ignore the feature combinations or structural information in a slightly larger range, thereby reducing the recognition accuracy.

[0050] According to the invoice text information recognition method of the present invention, by dynamically generating the weights of the convolutional kernel, the optimized YOLOv8 model can adaptively adjust the convolutional operation according to the input, thereby enhancing the expression ability of the model and improving the text detection ability. In addition, by replacing the ordinary convolution in the detection head of the YOLOv8 model with partial convolution (PConv), the amount of computation and the number of parameters are significantly reduced, thereby reducing the complexity of the model and improving the text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction, and utilizes the effectiveness of STN in geometric correction to effectively alleviate geometric deformations such as text distortion and tilt in the input image, thereby improving the text recognition efficiency. By increasing the proportion of the local mixing in the SVTR to enhance the local feature perception ability and text information capture ability of the SVTR, making it more suitable for scenarios with relatively fixed formats but complex local details such as invoices, so as to more effectively recognize various structured fields on the invoice.

[0051] Embodiment 2 This embodiment provides an invoice text information recognition system, in which the invoice text information is recognized by using the above recognition method. The system includes: An acquisition module, configured to acquire an invoice image as a detection data set and label the key text in the invoice image to obtain training samples; A first construction module, configured to construct a YOLOv8 model based on the dynamic convolution of ParameterNet. The YOLOv8 model includes a backbone network, a neck network, and multiple detection heads. The dynamic convolution in ParameterNet is introduced into the C2f modules in the backbone network and the neck network to enhance the feature extraction ability of the YOLOv8 model; The training output module is used to introduce a partial convolution strategy in each detection head to reduce the computational complexity of the model's detection head, improve the inference speed and resource utilization efficiency of the model, and enhance the text detection ability. The YOLOv8 model is trained according to the training samples to output the text regions of the training samples; The second construction module is used to construct an invoice text information recognition dataset according to the text regions; The third construction module is used to construct a text recognition model. The text recognition model includes a spatial transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model. The scene text recognition network based on a single vision model includes local mixing and global mixing; The shape of the recognition dataset is corrected according to the spatial transformation network, and the local feature perception ability and text information capture ability of the SVTR are enhanced by increasing the proportion of the number of local mixing in the SVTR; The output module is used to input the invoice text information recognition dataset into the text recognition model to output structured invoice text information. An invoice text information recognition system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.

[0052] An invoice text information recognition system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0053] An invoice text information recognition system provided by an embodiment of the present application can implement Figure 1 each process implemented by a method embodiment of an invoice text information recognition method. To avoid repetition, it will not be elaborated here.

[0054] According to the invoice text information recognition system of the present invention, by dynamically generating convolutional kernel weights, the optimized YOLOv8 model can adaptively adjust convolutional operations according to the input, thereby enhancing the model's expressive ability and improving text detection ability. In addition, by replacing the ordinary convolution in the YOLOv8 detection head with partial convolution (PConv), the amount of computation and the number of parameters are significantly reduced, thus reducing the complexity of the model and improving text detection ability. The invoice text information recognition method of the present invention integrates STN based on TPS at the input end of SVTR for text image correction, and utilizes the effectiveness of STN in geometric correction to effectively alleviate geometric deformations such as text distortion and tilt existing in the input image, thereby improving text recognition efficiency. By increasing the proportion of the local mixture in SVTR to enhance the local feature perception ability and text information capture ability of SVTR, making it more suitable for scenarios with relatively fixed layouts but complex local details such as invoices, so as to more effectively identify various structured fields on the invoice.

[0055] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the invoice text information recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0056] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned embodiment of the invoice text information recognition method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0057] Wherein, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc.

[0058] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the invention.

[0059] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example.

[0060] Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The mention of "embodiment" in this text means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0061] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for identifying invoice text information, characterized in that: include: Obtain invoice images as a detection dataset, and annotate key texts in the invoice images to obtain training samples; Constructing a YOLOv8 model based on dynamic convolution of ParameterNet, wherein the YOLOv8 model includes a backbone network, a neck network, and multiple detection heads, and introducing the dynamic convolution in the ParameterNet into the C2f module in the backbone network and the neck network to enhance the feature extraction capability of the YOLOv8 model; Introducing a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and training the YOLOv8 model according to the training sample to output the text region of the training sample; Constructing an invoice text information recognition data set based on the text area; Constructing a text recognition model, the text recognition model includes a space transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model, the scene text recognition network based on the single vision model includes local mixing and global mixing; correcting the shape of the recognition data set according to the space transformation network, and enhancing the local feature perception ability and text information capture ability of the scene text recognition network based on the single vision model by increasing the proportion of the local mixing in the scene text recognition network based on the single vision model; The invoice text information recognition data set is input into the text recognition model to output structured invoice text information.

2. The invoice text information recognition method according to claim 1, characterized in that: The spatial transformation network includes a positioning network, a grid generator and a sampler, wherein the positioning network is a convolutional neural network; Correcting the shape of the recognition data set according to the spatial transformation network specifically includes: An input image in an invoice text information recognition dataset is received using a localization network, and is used to predict the positions of a set of key control points in the input image, wherein the key control points represent areas in the image that need to be deformed; Using a mesh generator to generate a deformed mesh according to the positions of the key control points using a thin plate spline interpolation algorithm, and defining a new position of each pixel in the input image in the output corrected image according to the deformed mesh; The input image is sampled using a sampler according to the deformed grid to generate an output rectified image.

3. The invoice text information recognition method according to claim 1, characterized in that: Increasing the proportion of the local mixture in the scene text recognition network based on the single vision model to enhance the local feature perception ability and text information capture ability of the scene text recognition network based on the single vision model specifically includes: Receiving the output rectified image to the image block embedding module for partitioning to obtain image blocks of fixed size; Performing convolution mapping according to the fixed-size image block to obtain an initial feature sequence of an output rectified image; Passing the initial feature sequence through a first mixing module and a second mixing module to generate a compressed feature sequence, wherein the first mixing module and the second mixing module both include a plurality of local mixings; Passing the compressed feature sequence through a third mixing module to generate a final feature sequence, wherein the third mixing module includes one local mixing and five global mixings; The final feature sequence is passed through a fully connected layer for character classification to output a recognition result.

4. The invoice text information recognition method according to claim 1, characterized in that: The dynamic convolutional network based on ParameterNet is a network generated by combining the convolution kernels and weights of multiple dynamic experts, and using multiple predefined convolution kernels. and the dynamic coefficients corresponding to the convolution kernel The weighted combination of is generated; Output feature map of the YOLOv8 model based on dynamic convolutional network The calculation formula is: in, Represents the input feature map, and its dimension is expressed as , represents the number of input channels, H represents the input height, and W represents the input width; Represents the output feature map, which is the output result of the dynamic convolution operation, and its dimension is expressed as , Indicates the number of output channels, Indicates the output height, Indicates the output width, represents the i-th dynamic expert, M represents the number of dynamic experts; Dynamic coefficient It is generated based on the input features through global average pooling and multi-layer perceptron. The formula is: Pool means global average pooling, and MLP means multi-layer perceptron.

5. The invoice text information recognition method according to claim 1, characterized in that: During the annotation process, a rectangular bounding box is used to select the key text of each invoice image to obtain the corresponding target label file, thereby obtaining a training sample.

6. The invoice text information recognition method according to claim 1, characterized in that: In the feature extraction stage of the backbone network, the training samples are subjected to multiple downsampling convolution operations to extract multi-scale features, dynamic convolutions are interspersed under multiple downsampling convolutions in the backbone network to extract features of different scales, and finally enter the SPPF module in the YOLOv8 model to output feature results, completing the feature extraction process of the backbone network.

7. The invoice text information recognition method according to claim 1, characterized in that: In the feature fusion stage of the neck network, it includes: Perform upsampling convolution on the feature results output by the backbone network, and perform splicing operation with the different scale features output by the dynamic convolution in the backbone network to obtain the spliced ​​scale features; The spliced ​​features of different scales are fused according to dynamic convolution to form a multi-scale semantic information flow.

8. An invoice text information recognition system, characterized in that: The invoice text information is identified by using the identification method according to any one of claims 1 to 7, wherein the system comprises: An acquisition module is used to acquire invoice images as a detection data set and annotate key texts in the invoice images to obtain training samples; A first construction module is used to construct a YOLOv8 model based on dynamic convolution of ParameterNet, wherein the YOLOv8 model includes a backbone network, a neck network, and multiple detection heads, and the dynamic convolution in the ParameterNet is introduced into the C2f module in the backbone network and the neck network to enhance the feature extraction capability of the YOLOv8 model; A training output module, used for introducing a partial convolution strategy in each detection head to reduce the computational complexity of the model detection head, and training the YOLOv8 model according to the training sample to output the text area of ​​the training sample; A second construction module is used to construct an invoice text information recognition data set according to the text area; A third construction module is used to construct a text recognition model, wherein the text recognition model includes a space transformation network based on thin plate spline interpolation and a scene text recognition network based on a single vision model, wherein the scene text recognition network based on the single vision model includes local mixing and global mixing; according to the space transformation network, the shape of the recognition data set is corrected, and the local feature perception ability and text information capture ability of the scene text recognition network based on the single vision model are enhanced by increasing the proportion of the local mixing in the scene text recognition network based on the single vision model; The output module is used to input the invoice text information recognition data set into the text recognition model to output structured invoice text information.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the invoice text information recognition method as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the invoice text information recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target detection method based on partial convolution embedding and aggregation distribution mechanism

    CN117671414A

  • Bill information identification method, device and equipment based on OCR (Optical Character Recognition) and storage medium

    CN118397642A

  • Value-added tax invoice image content segmentation method and system based on YOLOv8-Seg

    CN119314194A

  • High-altitude electric power operation violation identification method

    CN119785431A

  • Arithmetic question marking system based on mixnet-yolov3 and convolutional recurrent neural network (CRNN)

    WO2022147965A1