Bill text detection method based on attention mechanism and feature enhancement
By introducing the attention mechanism and feature enhancement method in bill detection, and utilizing the multi-scale residual module and Transformer block, the adhesion problem of text detection in complex backgrounds is solved, and efficient segmentation and positioning of low-quality text is achieved.
Patent Information
- Application Number
- CN202510583980.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing bill detection algorithms have difficulty effectively processing diverse texts in complex backgrounds, especially curved, tilted or handwritten texts, and their computational efficiency is low. Traditional methods have mismatching problems when detecting contiguous texts.
A method based on attention mechanism and feature enhancement is adopted to extract structural information features through a multi-scale residual module, and Transformer is introduced in the neck network for global feature modeling. Combined with the adaptive feature fusion module, the correlation between text instances is captured and a binary image is generated to extract text areas.
It improves the ability to handle adhesion problems in text areas, accurately segments independent text instances, enhances the detection performance of low-quality text, and improves the ability to locate text instances.
Smart Images

Figure CN120635925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text detection technology, and in particular, to a bill text detection method based on attention mechanism and feature enhancement. Background Art
[0002] As a key task in computer vision, scene text detection has evolved through both traditional visual methods and deep learning techniques, gradually forming a multimodal, high-precision technology system. Traditional methods rely on underlying image features to achieve text localization through two main approaches: structural feature analysis and candidate region generation. Among these, the maximum stable extremum region and stroke width transformation methods extract connected domains through geometric features. These methods, combined with sliding windows and cascade classifiers to improve detection efficiency, continue to have practical value in embedded devices.
[0003] Deep learning technology has driven rapid development in this field, forming three major paradigms: regression-based detection frameworks solve multi-directional text detection by improving the anchor mechanism (such as the quadrilateral anchor box of TextBoxes++); segmentation-based paradigms (such as PSENet and Mask TextSpotter) convert detection into pixel-level classification, effectively processing text of arbitrary shapes; end-to-end systems integrate detection and recognition modules and combine them with the Transformer architecture (such as DETR) to achieve global context modeling.
[0004] There are some problems with the existing bill detection algorithms, as follows:
[0005] Traditional methods mainly rely on image processing technology and manual feature extraction, including connected region-based methods, sliding window-based methods, and edge detection-based methods. The connected region-based method screens out candidate areas that may contain text by segmenting the connected regions in the image; the sliding window-based method uses a multi-scale window to scan the image and classifies the text area using texture features; the edge detection-based method locates the text by capturing the density and directional consistency of the text edge. These methods have certain effects in scenes with clear text and regular layouts, but it is difficult to represent the diversity of text in complex backgrounds. The rule logic cannot adapt to the changing layout structure, and the computational efficiency is low. The details are as follows:
[0006] While segmentation-based deep models (such as PSENet, DBNet, LSAE, CRAFT, SAST, DBNet++, etc.) can generate pixel-level predictions, they may perform poorly for some non-standard text (such as curved, tilted, or handwritten text). In particular, when text is tilted at large angles, deformed, or extremely curved, traditional segmentation-based methods may not be able to capture the boundaries of the text region well.
[0007] Anchor-based detectors (such as TextBoxes, TextBoxes++, Textsnake, and FCENet) use pre-defined horizontal and multi-directional rectangular boxes, making it difficult to fit the irregular contours of contiguous text. Curved text detectors (such as ABCNet), while incorporating Bezier curve parameterization, still suffer from control point mismatching issues when dealing with overlapping text (e.g., multi-layered text on billboards). Summary of the Invention
[0008] In response to the aforementioned technical issues, a method for invoice text detection based on an attention mechanism and feature enhancement is provided. By introducing a multi-scale residual module into the backbone network, this invention effectively extracts structural information features and enhances the network's feature extraction capabilities. Furthermore, by introducing a feature enhancement module into the neck network, the model enhances the extraction of global semantic information and captures correlations between different text instances, accurately detecting text regions within medical invoices.
[0009] The technical means adopted in the present invention are as follows:
[0010] A bill text detection method based on attention mechanism and feature enhancement, including:
[0011] S1. Obtain a medical bill image as the input image x to be detected;
[0012] S2. Build a text detection network, input the input image x to be detected into the built text detection network, and output a binary image B;
[0013] S3. Extract the outline of the text area from the binary image B through edge detection to obtain a detection frame of the text area.
[0014] Furthermore, step S2 specifically includes:
[0015] S21. Build the backbone network of the text detection network, perform feature extraction based on multi-scale residual blocks, and output feature maps C1, C2, C3, and C4 at four different scales.
[0016] S22. Build the neck network part of the text detection network, perform global feature modeling based on Transformer, and output feature maps P1, P2, P3, and P4 of four different scales corresponding to feature maps C1, C2, C3, and C4;
[0017] S23, constructing an adaptive feature fusion module of the text detection network, performing adaptive multi-scale feature fusion, and obtaining a weighted fusion feature map F;
[0018] S24. Construct a text region labeling module of the text detection network to perform text region labeling. Use the fused feature map F as the input feature map, derive two branches respectively, obtain the probability map and threshold map through convolution and upsampling operations, and then obtain the binarized map B through differentiable binarization operations.
[0019] Furthermore, in step S21, the ResNet50 backbone network is improved by replacing the original residual block containing 3×3 convolution with a multi-scale residual block to enhance feature extraction and modeling capabilities, specifically including:
[0020] S211: Input bill image After a 7×7 convolution kernel, a convolution layer with a stride of 2, and a 3×3 maximum pooling layer with a stride of 2, the output feature map has a size of
[0021] S212, in the first residual stage of the backbone network, contains 3 multi-scale residual blocks, the input is the feature map output in step S211, the size is The output feature map is C1, and its size is
[0022] S213, in the second residual stage of the backbone network, it contains 4 multi-scale residual blocks, and the input is the feature map C1 output in step S212, with a size of The output feature map is C2, and the size is
[0023] S214, in the third residual stage of the backbone network, contains 6 multi-scale residual blocks, and the input is the feature map C2 output in step S213, with a size of The output feature map is C3, and the size is
[0024] S215, in the fourth residual stage of the backbone network, it contains 3 multi-scale residual blocks, and the input is the feature map C3 output in step S214, with a size of The output feature map is C4, and the size is
[0025] Furthermore, the multi-scale residual block includes a first 1×1 convolution, a layered convolution module, a pyramid pooling module and a second 1×1 convolution, wherein:
[0026] First, the dimensionality is reduced through the first 1×1 convolution; then, the layered convolution module is used to enhance features at different scales, which can enhance the multi-scale feature representation capability at a finer level; then, the pyramid pooling module uses two-size average pooling to aggregate global and local context, calculates feature weights at different scales, and adaptively adjusts attention weights to make the model more focused on the text area. At the same time, convolution is used to learn the relationship between channels and improve the performance of channel attention; finally, the dimensionality is increased through the second 1×1 convolution.
[0027] Furthermore, step S22 specifically includes:
[0028] S221, applying convolution blocks to the four feature maps C1, C2, C3, and C4 of different scales outputted from the feature extraction stage to unify the number of channels;
[0029] S222. Use the feature enhancement module to extract global semantic information and spatial position information to capture the relationship between different text instances, and output feature maps P1, P2, P3 and P4. The size of each output feature map is the same as the size of its corresponding input feature map.
[0030] Furthermore, the feature enhancement module is composed of N improved Transformer blocks repeatedly stacked, with the number of Transformer blocks from deep features to shallow features being 2, 3, 3, and 4 respectively; the Transformer block consists of a positional encoder, a dynamic sparse attention module, a feedforward network, a ReLU activation function, layer normalization, and a residual connection, wherein:
[0031] Position encoding uses 3×3 depth convolution to implicitly encode relative position information, and the output dimension remains unchanged;
[0032] The number of attention heads is 8. The dynamic sparse attention module and layer normalization are used to encode the local and fine-grained information of the text on the input feature map, while the output dimension remains unchanged.
[0033] After residual connection, the features are input into the feedforward network and layer normalization. The feedforward network consists of two linear layers and the output dimension remains unchanged.
[0034] Furthermore, step S23 specifically includes:
[0035] S231, performing a splicing operation on the scaled feature maps;
[0036] S232. Obtain intermediate features through convolution operations, and apply spatial attention mechanisms to the obtained intermediate features to obtain attention weights, so that the network can learn the importance of each position;
[0037] S233. Multiply the obtained weight by the corresponding intermediate feature to obtain a weighted fusion feature map F.
[0038] Furthermore, step S24 specifically includes:
[0039] S241, using the fused feature map F as the input feature map of the text region localization stage, and performing convolution, upsampling, and Sigmoid activation operations on the fused feature map F to generate a probability map M and a threshold map T of the same size as the original image;
[0040] S242, based on the generated probability map M and threshold map T, generate a binary map B through a differentiable binarization (DB) operation;
[0041] S243. Design the network loss function. During the training process, the number of training samples is N, and the loss function includes the probability map loss L. m , Binarization image loss L b and threshold map loss L t , the loss function is defined as the weighted sum of three loss functions, and the calculation formula is as follows:
[0042] L=L m +ω×L b +σ×L t
[0043] Among them, the weight coefficients ω and σ are set to 1 and 10 respectively; L m and L b BCE loss is usually used, L t Using L1 loss, the calculation formula is as follows:
[0044]
[0045] in, and are the actual label values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample, respectively. and are the predicted values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample respectively.
[0046] Furthermore, step S3 specifically includes:
[0047] S31, extracting the outline of the text area from the binary image through edge detection;
[0048] S32. For each contour, a contour of a text area with fewer vertices is returned by a polygonal approximation method, and finally a detection result of text detection is obtained.
[0049] Compared with the prior art, the present invention has the following advantages:
[0050] 1. The present invention provides a method for detecting invoice text based on an attention mechanism and feature enhancement. This method effectively addresses the problem of overlapping text regions and accurately segments independent text instances. Furthermore, the present method demonstrates excellent detection performance even when low-quality text is missed in medical invoices.
[0051] 2. The present invention provides a bill text detection method based on attention mechanism and feature enhancement. After adding multi-scale residual blocks and Transformer blocks, the ability to locate text instances is improved.
[0052] Based on the above reasons, the present invention can be widely promoted in fields such as text detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0054] Figure 1 Flow chart of the method of the present invention.
[0055] Figure 2 This is the backbone network structure diagram of the text detection network of the present invention.
[0056] Figure 3 This is a network structure diagram of the multi-scale residual block of the present invention.
[0057] Figure 4 This is the network structure diagram of the layered convolution module of the present invention.
[0058] Figure 5 This is the neck network structure diagram of the text detection network of the present invention.
[0059] Figure 6 This is the network structure diagram of the Transformer block of the present invention.
[0060] Figure 7 This is a network structure diagram of the adaptive feature fusion module of the present invention.
[0061] Figure 8 This is a network structure diagram of the text area marking module of the present invention.
[0062] Figure 9 This is a binarized image provided by an embodiment of the present invention.
[0063] Figure 10This is a result diagram of the detection method provided by an embodiment of the present invention.
[0064] Figure 11 This is a diagram of the DBNet++ detection results provided by an embodiment of the present invention.
[0065] Figure 12 This is a diagram of the improved DBNet++ detection results provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0066] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0067] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.
[0068] like Figure 1 As shown, the present invention provides a bill text detection method based on attention mechanism and feature enhancement, which aims to efficiently detect text areas in bill images, including:
[0069] S1. Obtain a medical bill image as the input image x to be detected;
[0070] S2. Build a text detection network, input the input image x to be detected into the built text detection network, and output a binary image B;
[0071] S3. Extract the outline of the text area from the binary image B through edge detection to obtain a detection frame of the text area.
[0072] When implementing the invention, please refer to the preferred embodiment of the invention. Figure 1 Step S2 specifically includes:
[0073] S21. Build the backbone network of the text detection network, perform feature extraction based on multi-scale residual blocks, and output feature maps C1, C2, C3, and C4 at four different scales.
[0074] S22. Build the neck network part of the text detection network, perform global feature modeling based on Transformer, and output feature maps P1, P2, P3, and P4 of four different scales corresponding to feature maps C1, C2, C3, and C4;
[0075] S23, constructing an adaptive feature fusion module of the text detection network, performing adaptive multi-scale feature fusion, and obtaining a weighted fusion feature map F;
[0076] S24. Construct a text region labeling module of the text detection network to perform text region labeling. Use the fused feature map F as the input feature map, derive two branches respectively, obtain the probability map and threshold map through convolution and upsampling operations, and then obtain the binarized map B through differentiable binarization operations.
[0077] In specific implementation, as a preferred embodiment of the present invention, in step S21, the ResNet50 backbone network is improved, and the original residual block containing 3×3 convolution is replaced by a multi-scale residual block to improve feature extraction and modeling capabilities. The overall structure is as follows Figure 2 As shown, it can effectively handle the problem of missed detection of low-contrast text in medical bills, including:
[0078] S211: Input bill image After a 7×7 convolution kernel, a convolution layer with a stride of 2, and a 3×3 maximum pooling layer with a stride of 2, the output feature map has a size of
[0079] S212, in the first residual stage of the backbone network, contains 3 multi-scale residual blocks, the input is the feature map output in step S211, the size is The output feature map is C1, and its size is
[0080] S213, in the second residual stage of the backbone network, it contains 4 multi-scale residual blocks, and the input is the feature map C1 output in step S212, with a size of The output feature map is C2, and the size is
[0081] S214, in the third residual stage of the backbone network, contains 6 multi-scale residual blocks, and the input is the feature map C2 output in step S213, with a size of The output feature map is C3, and the size is
[0082] S215, in the fourth residual stage of the backbone network, it contains 3 multi-scale residual blocks, and the input is the feature map C3 output in step S214, with a size of The output feature map is C4, and the size is
[0083] In this embodiment, the ResNet50 backbone network is improved during the feature extraction phase, employing multi-scale residual blocks instead of the original residual blocks containing 3×3 convolutions to enhance feature extraction and modeling capabilities. The input is a bill text image x. Through the synergistic effect of layered convolution and pyramid pooling modules, the network is able to adaptively and dynamically calibrate channel weights, capturing more fine-grained information during multi-scale feature extraction and effectively fusing contextual information at different levels to obtain richer feature expressions and structural information. After feature extraction by the backbone network, four feature maps C1, C2, C3, and C4 at different scales are output. The backbone network extracts multi-level spatial feature representations, providing a richer hierarchical feature representation for subsequent feature screening and fusion.
[0084] In a specific implementation, as a preferred embodiment of the present invention, the multi-scale residual block includes a first 1×1 convolution, a layered convolution module, a pyramid pooling module, and a second 1×1 convolution, wherein:
[0085] First, the dimensionality is reduced through the first 1×1 convolution; then, the layered convolution module is used to enhance features at different scales, which can enhance the multi-scale feature representation capability at a finer level; then, the pyramid pooling module uses two-size average pooling to aggregate global and local context, calculates feature weights at different scales, and adaptively adjusts attention weights to make the model more focused on the text area. At the same time, convolution is used to learn the relationship between channels and improve the performance of channel attention; finally, the dimensionality is increased through the second 1×1 convolution.
[0086] In this embodiment, the network structure diagram of the multi-scale residual block is as follows: Figure 3 As shown, the first multi-scale residual block of the second residual stage is used as an example. The first input is The feature map of is reduced by a convolution layer with a 1×1 convolution kernel and a step size of 2, and the output feature map is Then through the layered convolution module, the layered convolution module is implemented by three operations: Split, Conv and Concat. The overall network architecture Figure 4 As shown, the Split operation is to divide the channel dimension into 4 feature map subsets, and for each subset Apply 3×3 convolution and batch normalization operations, and the output feature map size remains unchanged. Then, apply hierarchical residual convolution on each subset. Different groups of convolution operators are fused in the form of hierarchical residual connections to increase the model's ability to express features of different scales. Finally, all feature maps are spliced together, and the output feature map is The feature map is then input into the pyramid attention module. The input feature map is It is divided into 4 feature map subsets, each with 32 channels. Each feature map subset is passed through 1×1 and 3×3 pooling layers respectively, and then the attention weights are generated through 1×1 convolution layer and Sigmoid activation function after splicing. The weight size is 1×1×32. Then the attention weights of the 4 feature map subsets are concat spliced. The size of the spliced weight is 1×1×128. The weights are combined with the input feature map. Multiply and output the weighted feature map, the size is Finally, the dimension is increased by 1×1 convolution, and the output feature map is
[0087] In specific implementation, as a preferred embodiment of the present invention, step S22 specifically includes:
[0088] S221, applying convolution blocks to the four feature maps C1, C2, C3, and C4 of different scales outputted from the feature extraction stage to unify the number of channels;
[0089] S222. Use the feature enhancement module to extract global semantic information and spatial position information to capture the relationship between different text instances, and output feature maps P1, P2, P3 and P4. The size of each output feature map is the same as the size of its corresponding input feature map.
[0090] In specific implementation, as a preferred embodiment of the present invention, the feature enhancement module is composed of N improved Transformer blocks repeatedly stacked, and the overall structure of the module is as follows: Figure 5 As shown in , the number of Transformer blocks from deep features to shallow features are 2, 3, 3, and 4 respectively; the network structure of the Transformer block is as follows Figure 6 As shown in Figure 1, the Transformer block consists of a positional encoder, a dynamic sparse attention module (BRT), a feedforward network, a ReLU activation function, layer normalization, and a residual connection, where:
[0091] Position encoding uses 3×3 depth convolution to implicitly encode relative position information, and the output dimension remains unchanged;
[0092] The number of attention heads is 8. The dynamic sparse attention module and layer normalization are used to encode the local and fine-grained information of the text on the input feature map, while the output dimension remains unchanged.
[0093] After residual connection, the features are input into the feedforward network and layer normalization. The feedforward network consists of two linear layers and the output dimension remains unchanged.
[0094] In this embodiment, the feature enhancement module further optimizes the multi-scale features extracted by the backbone network. The input feature maps are C1, C2, C3, and C4. By adding a dynamic sparse attention mechanism to the Transformer block stack, more global semantic information is captured. Through bottom-up feature extraction, top-down feature fusion, and lateral connections, features of different scales are effectively integrated, so that high-level features have richer semantic information and low-level features retain more accurate spatial location information, thereby improving the detection ability of text. The output is four feature maps of different scales P1, P2, P3, and P4 corresponding to the input feature maps.
[0095] Taking the input C2 feature map as an example, first C2 passes through the convolution block to unify the number of channels, and the output feature map is Y2. The size of the feature map is Then the output feature map T2 is obtained by the Transformer block. The size of the feature map remains unchanged. Then it is upsampled and added element by element to the feature map T1. The size of the output feature map is The feature map T1 is obtained by passing the feature map C1 through the convolution block and the Transformer block. Finally, the output feature map P1 is obtained after upsampling, and the size is
[0096] Taking the input Y2 feature map as an example, the input feature map of the dynamic sparse attention module (BRT) is the output feature map obtained by position encoding the feature map Y2. Figure X 2. This module first transforms the features Figure X 2 is divided into non-overlapping regions and each region is mapped into a query (Q) and a key-value pair (K, V) through a linear transformation module, as shown in the following formula:
[0097]
[0098] Then, the similarity between each region is calculated. Each element of the similarity matrix A represents the relationship between regions. Dot products are usually used to measure the similarity between regions, resulting in a "directed graph" representation between regions. Each region interacts with other regions based on the similarity matrix. The formula for calculating similarity is as follows:
[0099] A r =Q r (K r ) T
[0100] Find other regions with high similarity in the similarity matrix, and only take the top k regions as relevant regions to participate in fine-grained operations, as shown in the following formula:
[0101] I r=topIndex(A r )
[0102] Among them, I r Denotes the k most relevant areas of the ith zone in the ith row, and obtains the area-to-area routing index matrix I r .
[0103] After obtaining the region-to-region index matrix I r Afterwards, attention is applied to the tokens within each region to generate a feature map O2. This is achieved by calculating the attention relationship between the tokens within the region, enabling fine-grained information flow at the token level. Since these routing regions are expected to be scattered across the entire feature map, and modern GPUs rely on merged memory operations that load blocks of dozens of consecutive bytes at once, it is necessary to aggregate the key and value vectors. This process is shown in the following equation:
[0104] K g =gather(K r ,I r ), V g =gather(V r ,I r )
[0105] O2=Attention(Q r ,K g ,V g )+LCE(V r ).
[0106] In specific implementation, as a preferred embodiment of the present invention, step S23 specifically includes:
[0107] S231, performing a splicing operation on the scaled feature maps;
[0108] S232. Obtain intermediate features through convolution operations, and apply spatial attention mechanisms to the obtained intermediate features to obtain attention weights, so that the network can learn the importance of each position;
[0109] S233. Multiply the obtained weight by the corresponding intermediate feature to obtain a weighted fusion feature map F.
[0110] In this embodiment, the adaptive feature fusion module does not use a simple addition method to better fuse features of different scales. Instead, it allows the network to select the importance of features of different scales and locations and dynamically aggregate the features. A more specific implementation is as follows:
[0111] The four feature maps P1, P2, P3 and P4 output from the global feature modeling stage are used as the input of the adaptive multi-scale feature fusion stage. The sizes of the four feature maps are That is, the number of channels C is 64, and the width and height are and Then, after the adaptive multi-scale feature fusion module, the overall structure of the adaptive multi-scale feature fusion module is as follows: Figure 7 As shown, the input feature map is first spliced along the channel, and the output feature map size is Then the intermediate feature S is obtained through a 3×3 convolution layer, and its size is The intermediate feature S passes through an attention module to obtain the attention weight A, whose size is The attention module is composed of two convolutions and Sigmoid, and the generated weights are used to highlight specific areas. The weight A is then replicated into 64 copies along the channel dimension through a broadcast mechanism, which is the same size as the input feature map after splicing. That is, the size of weight A is Finally, weighted multiplication is performed to obtain the final fusion feature F, the size of which is
[0112] In specific implementation, as a preferred embodiment of the present invention, step S24 specifically includes:
[0113] S241, the fused feature map F is used as the input feature map of the text area positioning stage, and the fused feature map F is subjected to convolution, upsampling and Sigmoid activation operations to generate a probability map M and a threshold map T with the same size as the original image; in this embodiment, the specific process is as follows Figure 8 As shown in the figure, convolution and upsampling are used to adjust the number of channels in the feature map. The last layer that generates the probability map M uses Sigmoid activation, and the output value is between 0 and 1, indicating the probability of each pixel belonging to the text area. The last layer that generates the threshold map T uses Sigmoid activation, and the output value is between 0 and 1, indicating the threshold of each pixel. It is used to compare with the probability value of the probability map T to determine whether the pixel belongs to the text area.
[0114] S242, based on the generated probability map M and threshold map T, generate the following through differentiable binarization (DB) operation: Figure 9 The binary image B shown in the figure; in this embodiment, the DB operation process has a different threshold for each pixel point. The threshold map is T, which can be learned from the network. Where k is the amplification factor, set k = 50. The formula of the step function is shown as follows:
[0115]
[0116] Among them, (i, j) represents the coordinate position in the feature map.
[0117] S243. Design the network loss function. During the training process, the number of training samples is N, and the loss function includes the probability map loss L. m , Binarization image loss L b and threshold map loss L t , the loss function is defined as the weighted sum of three loss functions, and the calculation formula is as follows:
[0118] L=L m +ω×L b +σ×L t
[0119] Among them, the weight coefficients ω and σ are set to 1 and 10 respectively; L m and L b BCE loss is usually used, L t Using L1 loss, the calculation formula is as follows:
[0120]
[0121] in, and are the actual label values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample, respectively. and are the predicted values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample respectively.
[0122] In specific implementation, as a preferred embodiment of the present invention, step S3 specifically includes:
[0123] S31. In this embodiment, after obtaining the binary image, the pixels in the text area in the binary image are set to 1, and the pixels in the non-text area are set to 0. Then, the outline of the text area is extracted from the binary image through edge detection.
[0124] S32, for each contour, a contour of a text area with fewer vertices is returned by a polygonal approximation method, and finally the detection result of the text detection is obtained, such as Figure 10 As shown in the figure. In the binarized image, pixels with probability scores greater than the threshold are considered text areas, while pixels with probability scores less than the threshold are considered background areas. This algorithm can effectively address the problem of text region adhesion and accurately segment independent text instance regions. Furthermore, the text detection method proposed in this paper also demonstrates excellent detection performance for low-quality text in medical bills.
[0125] exist Figure 11 and Figure 12The comparison chart of the results obtained using different detection algorithms is shown in Figure 2. As can be seen from the figure, the original DBNet++ lacks the ability to detect edge pixels of text instances, resulting in inaccurate detection of character adhesion. The addition of multi-scale residual blocks and Transformer blocks improves the ability to locate text instances.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A bill text detection method based on attention mechanism and feature enhancement, characterized in that: include: S1. Obtain a medical bill image as the input image x to be detected; S2. Build a text detection network, input the input image x to be detected into the built text detection network, and output a binary image B; S3. Extract the outline of the text area from the binary image B through edge detection to obtain a detection frame of the text area.
2. The method for bill text detection based on attention mechanism and feature enhancement according to claim 1, characterized in that: Step S2 specifically includes: S21. Build the backbone network of the text detection network, perform feature extraction based on multi-scale residual blocks, and output feature maps C1, C2, C3, and C4 at four different scales. S22. Build the neck network part of the text detection network, perform global feature modeling based on Transformer, and output feature maps P1, P2, P3, and P4 of four different scales corresponding to feature maps C1, C2, C3, and C4; S23, constructing an adaptive feature fusion module of the text detection network, performing adaptive multi-scale feature fusion, and obtaining a weighted fusion feature map F; S24. Construct a text region labeling module of the text detection network to perform text region labeling. Use the fused feature map F as the input feature map, derive two branches respectively, obtain the probability map and threshold map through convolution and upsampling operations, and then obtain the binarized map B through differentiable binarization operations.
3. The method for bill text detection based on attention mechanism and feature enhancement according to claim 2, characterized in that: In step S21, the ResNet50 backbone network is improved by replacing the original residual block containing 3×3 convolution with a multi-scale residual block to enhance feature extraction and modeling capabilities. Specifically, the following steps are performed: S211: Input bill image After a 7×7 convolution kernel, a convolution layer with a stride of 2, and a 3×3 maximum pooling layer with a stride of 2, the output feature map has a size of S212, in the first residual stage of the backbone network, contains 3 multi-scale residual blocks, the input is the feature map output in step S211, the size is The output feature map is C1, and its size is S213, in the second residual stage of the backbone network, it contains 4 multi-scale residual blocks, and the input is the feature map C1 output in step S212, with a size of The output feature map is C2, and the size is S214, in the third residual stage of the backbone network, contains 6 multi-scale residual blocks, and the input is the feature map C2 output in step S213, with a size of The output feature map is C3, and the size is S215, in the fourth residual stage of the backbone network, it contains 3 multi-scale residual blocks, and the input is the feature map C3 output in step S214, with a size of The output feature map is C4, and the size is 4. The method for bill text detection based on attention mechanism and feature enhancement according to claim 3 is characterized in that: The multi-scale residual block includes a first 1×1 convolution, a layered convolution module, a pyramid pooling module, and a second 1×1 convolution, wherein: First, the dimensionality is reduced through the first 1×1 convolution; then, the layered convolution module is used to enhance features at different scales, which can enhance the multi-scale feature representation capability at a finer level; then, the pyramid pooling module uses two-size average pooling to aggregate global and local context, calculates feature weights at different scales, and adaptively adjusts attention weights to make the model more focused on the text area. At the same time, convolution is used to learn the relationship between channels and improve the performance of channel attention; finally, the dimensionality is increased through the second 1×1 convolution.
5. The method for bill text detection based on attention mechanism and feature enhancement according to claim 2, characterized in that: Step S22 specifically includes: S221, applying convolution blocks to the four feature maps C1, C2, C3, and C4 of different scales outputted from the feature extraction stage to unify the number of channels; S222. Use the feature enhancement module to extract global semantic information and spatial position information to capture the relationship between different text instances, and output feature maps P1, P2, P3 and P4. The size of each output feature map is the same as the size of its corresponding input feature map.
6. The method for bill text detection based on attention mechanism and feature enhancement according to claim 5, characterized in that: The feature enhancement module is composed of N improved Transformer blocks repeatedly stacked, with the number of Transformer blocks from deep features to shallow features being 2, 3, 3, and 4 respectively; the Transformer block consists of position encoding, dynamic sparse attention module, feedforward network, ReLU activation function, layer normalization, and residual connection, where: Position encoding uses 3×3 depth convolution to implicitly encode relative position information, and the output dimension remains unchanged; The number of attention heads is 8. The dynamic sparse attention module and layer normalization are used to encode the local and fine-grained information of the text on the input feature map, while the output dimension remains unchanged. After residual connection, the features are input into the feedforward network and layer normalization. The feedforward network consists of two linear layers and the output dimension remains unchanged.
7. The method for bill text detection based on attention mechanism and feature enhancement according to claim 2, characterized in that: Step S23 specifically includes: S231, performing a splicing operation on the scaled feature maps; S232. Obtain intermediate features through convolution operations, and apply spatial attention mechanisms to the obtained intermediate features to obtain attention weights, so that the network can learn the importance of each position; S233. Multiply the obtained weight by the corresponding intermediate feature to obtain a weighted fusion feature map F.
8. The method for bill text detection based on attention mechanism and feature enhancement according to claim 2, characterized in that: Step S24 specifically includes: S241, using the fused feature map F as the input feature map of the text region localization stage, and performing convolution, upsampling, and Sigmoid activation operations on the fused feature map F to generate a probability map M and a threshold map T of the same size as the original image; S242, based on the generated probability map M and threshold map T, generate a binary map B through a differentiable binarization (DB) operation; S243. Design the network loss function. During the training process, the number of training samples is N, and the loss function includes the probability map loss L. m , Binarization image loss L b and threshold map loss L t , the loss function is defined as the weighted sum of three loss functions, and the calculation formula is as follows: L=L m +ω×L b +σ×L t Among them, the weight coefficients ω and σ are set to 1 and 10 respectively; L m and L b BCE loss is usually used, L t Using L1 loss, the calculation formula is as follows: in, and are the actual label values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample, respectively. and are the predicted values of the i-th row and j-th column in the probability map, binary map, and threshold map of the S-th sample respectively.
9. The method for bill text detection based on attention mechanism and feature enhancement according to claim 1, characterized in that: Step S3 specifically includes: S31, extracting the outline of the text area from the binary image through edge detection; S32. For each contour, a contour of a text area with fewer vertices is returned by a polygonal approximation method, and finally a detection result of text detection is obtained.
Citation Information
Cited By
Medical document text detection method based on attention-guided multi-scale learning
CN121527782A