A scene text segmentation method, system, device and medium

By using SegFormer network and convolutional refinement module in scene text segmentation, combined with multi-layer perceptron and mixed loss function, the problems of insufficient accurate text segmentation in the existing technology, poor multi-directional text detection effect, and inconsistent text segmentation effects of different fonts, font sizes and colors are solved, and higher text segmentation accuracy is achieved.

CN117671692BActive Publication Date: 2025-05-13NANCHANG HANGKONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311699548.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-05-13
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

In the prior art, in the scene text segmentation, there are problems in the inaccurate small-target text segmentation, poor multi-direction text detection effect, and inconsistent text segmentation effects of different fonts, font sizes and colors.

Method used

The SegFormer network is used as the coarse-grained text segmentation module, and the fine-grained segmentation is combined with the convolutional refinement module. Multi-layer perceptrons are added to enhance supervised training of the model and mixed loss functions (including Focal loss, structural similarity index loss, and cross-border ratio loss) are used to improve segmentation accuracy.

Benefits of technology

Improves the accuracy of text segmentation, especially when dealing with small-target text and multi-directional text, and can better adapt text with different fonts, font sizes and colors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671692B_ABST
    Figure CN117671692B_ABST
Patent Text Reader

Abstract

The present invention discloses a scene text segmentation method, system, device and medium, and relates to the technical field of image segmentation. The method comprises: using a training set to train a scene text segmentation network to obtain a scene text segmentation model; the scene text segmentation network comprises a connected coarse-grained text segmentation module and a convolution refinement module, and the output of the convolution refinement module and the output of the coarse-grained text segmentation module are added pixel by pixel as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; during network training, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four prediction results output by the first multi-layer perceptron and the one prediction result output by the scene text segmentation network are used to supervise the training of the scene text segmentation network. The present invention improves the accuracy of text segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation, and in particular to a scene text segmentation method, system, device and medium. Background Art

[0002] In recent years, deep learning-based image segmentation technology has made rapid progress and has become one of the most popular research topics in computer vision. Compared to traditional segmentation algorithms, its segmentation accuracy is significantly improved and has been widely used in many fields, such as scene understanding, text detection and recognition, and autonomous driving. However, unlike the rapid development and application of segmentation technology in other fields, the application of scene text segmentation has lagged behind. This is partly due to the lack of a large and reliable dataset; partly due to the following problems with applying convolutional neural networks to scene text segmentation:

[0003] (1) The segmentation of small target text is not accurate enough. Due to the different sizes and shapes of scene text, the edges and details of some small target text are easily ignored by the convolution kernel in the convolution layer, resulting in inaccurate text segmentation results.

[0004] (2) Poor performance for multi-directional text detection. Since the convolution operation in convolutional neural networks (CNNs) is based on fixed convolution kernels, it cannot effectively handle multi-directional text segmentation tasks, such as tilted, curved, or distorted text.

[0005] (3) Variability in text appearance: The segmentation results for texts with different fonts, sizes, and colors are inconsistent. The training data for CNN models is usually large-scale text with a single font, size, and color. This makes the model's segmentation results for texts with different fonts, sizes, and colors less effective than originally expected. Summary of the Invention

[0006] The purpose of the present invention is to provide a scene text segmentation method, system, device and medium, which improve the accuracy of text segmentation.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A scene text segmentation method, comprising:

[0009] A scene text segmentation network is trained using a training set to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module and the output of the coarse-grained text segmentation module are added pixel by pixel as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used to perform fine-grained segmentation on the output of the coarse-grained text segmentation module;

[0010] Perform text segmentation on the image to be segmented using the scene text segmentation model to obtain a scene text segmentation result;

[0011] When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to perform supervised training on the scene text segmentation network.

[0012] Optionally, the convolution refinement module includes an encoder and a decoder, the output of the encoder is connected to the input of the decoder, the encoder includes a first convolution layer, a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence, the first convolution block, the second convolution block, the third convolution block and the fourth convolution all include convolution operations, downsampling operations, layer normalization operations and GELU activation function operations performed in sequence, the decoder includes a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block and a second convolution layer connected in sequence, the fifth convolution block, the sixth convolution block, the seventh convolution block and the eighth convolution block all include upsampling operations, convolution operations, layer normalization operations and GELU activation function operations performed in sequence; the first convolution layer is jump-connected to the second convolution layer, the first convolution block is jump-connected to the eighth convolution block, the second convolution block is jump-connected to the seventh convolution block, the third convolution block is jump-connected to the sixth convolution block, and the input of the first convolution layer and the output of the second convolution layer are added pixel by pixel as the output of the convolution refinement module.

[0013] Optionally, the four text segmentation feature maps output by the first multi-layer perceptron are respectively recorded as the first text segmentation feature map, the second text segmentation feature map, the second text segmentation feature map and the fourth text segmentation feature map, and the one text segmentation feature map output by the scene text segmentation network is recorded as the fifth text segmentation feature map;

[0014] The scene text segmentation network is trained using a hybrid loss function, which is expressed as:

[0015]

[0016] Among them, L represents the mixed loss, K is the number of text segmentation feature maps, K value is 5, α k The weight coefficient of the kth text segmentation feature map, l (k) is the loss value of the k-th text segmentation feature map, is the Focal loss of the k-th text segmentation feature map, is the structural similarity index loss of the k-th text segmentation feature map, is the intersection-over-union loss of the k-th text segmentation feature map.

[0017] Optionally, the Focal loss function for calculating the Focal loss is expressed as:

[0018] l focal =-α t (1-p t ) γ log(p t )

[0019] Among them, l focal Represents the intersection-over-union loss function value, p t represents the probability that the predicted output of the scene text segmentation network belongs to a positive sample, α t is the first weighting factor, and γ is the second weighting factor.

[0020] Optionally, the structural similarity index loss function for calculating the structural similarity index loss is expressed as:

[0021]

[0022] Among them, l ssim Represents the structural similarity index loss function value, μ x Represents the average value of the pixel value of the text segmentation feature map of the predicted output, σ x Represents the standard deviation of the pixel values ​​of the text segmentation feature map of the predicted output, μ y represents the average value of the pixel value of the real text segmentation map, σ y Represents the standard deviation of the pixel values ​​of the actual text segmentation map, σ xy It represents the covariance between the pixel value of the predicted text segmentation feature map and the pixel value of the actual text segmentation map, C1 represents the first constant, and C2 represents the second constant.

[0023] Optionally, the intersection-over-union loss function for calculating the intersection-over-union loss is expressed as:

[0024]

[0025] Among them, l iou Represents the intersection-over-union loss function value, H represents the height of the predicted output text segmentation feature map, W represents the width of the predicted output text segmentation feature map, P(r,c) represents the pixel value at position (r,c) on the predicted output text segmentation feature map, and G(r,c) represents the pixel value at position (r,c) on the actual text segmentation map.

[0026] Optionally, the SegFormer network includes a first Transformer encoder, a second Transformer encoder, a third Transformer encoder, a fourth Transformer encoder, MLP Layers, and a second multi-layer perceptron connected in sequence, wherein the outputs of the first Transformer encoder, the second Transformer encoder, and the third Transformer encoder are all connected to the input of the MLP Layers, and the first Transformer encoder is used to output The second Transformer encoder is used to output the feature map of the original image size. The third Transformer encoder is used to output the feature map of the original image size. The feature map of the original size, the fourth Transformer encoder is used to output A feature map of the original image size, where the original image size is the size of the image input to the first Transformer encoder;

[0027] The MLP Layers are used to Feature map of the original image size, Feature map of the original image size, Feature map of original image size and The feature maps of the original image size are subjected to dimensionality reduction operations respectively, and 4 text segmentation feature maps are output.

[0028] The present invention also discloses a scene text segmentation system, comprising:

[0029] A scene text segmentation model training module is used to train a scene text segmentation network using a training set to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module is added pixel by pixel to the output of the coarse-grained text segmentation module as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used to perform fine-grained segmentation on the output of the coarse-grained text segmentation module;

[0030] A scene text segmentation model application module is used to use the scene text segmentation model to perform text segmentation on the image to be segmented to obtain a scene text segmentation result;

[0031] When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to perform supervised training on the scene text segmentation network.

[0032] The present invention also discloses an electronic device, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the above-mentioned scene text segmentation method.

[0033] The present invention also discloses a computer-readable storage medium, characterized in that it stores a computer program, and the computer program is executed by a processor to implement the above-mentioned scene text segmentation method.

[0034] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0035] The present invention performs coarse-grained text segmentation on an image to be segmented through a SegFormer network, and then uses a convolution refinement module to perform fine-grained text segmentation on the output of the SegFormer network, thereby improving the accuracy of text segmentation. In addition, during the training process of a scene text segmentation network, a multi-layer perceptron (a first multi-layer perceptron) is added to the SegFormer network, the outputs of the MLP Layers in the SegFormer network are connected to the first multi-layer perceptron, and the scene text segmentation network is supervised and trained through four text segmentation feature maps output by the first multi-layer perceptron and one text segmentation feature map output by the scene text segmentation network, thereby further improving the accuracy of text segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] Figure 1 A schematic flow chart of a scene text segmentation method provided by an embodiment of the present invention;

[0038] Figure 2 A diagram showing the scene text segmentation network structure provided by an embodiment of the present invention;

[0039] Figure 3 A schematic diagram of the convolution refinement module results provided by an embodiment of the present invention;

[0040] Figure 4 A schematic diagram of a channel convolution operation provided by an embodiment of the present invention;

[0041] Figure 5 Schematic diagram of IOU segmentation provided by an embodiment of the present invention;

[0042] Figure 6 An example diagram of a data set provided by an embodiment of the present invention;

[0043] Figure 7 A schematic diagram of synthetic scene text data enhancement provided by an embodiment of the present invention;

[0044] Figure 8 Schematic diagram of segmentation results of different segmentation methods on different data sets provided by embodiments of the present invention;

[0045] Figure 9 Schematic diagram of IoU change graphs for different ablation experiment methods provided in an embodiment of the present invention;

[0046] Figure 10 Schematic diagram of F-Score changes in different ablation experiment methods provided by an embodiment of the present invention;

[0047] Figure 11 A schematic diagram of the segmentation results of an ablation experiment provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0049] The purpose of the present invention is to provide a scene text segmentation method, system, device and medium, which improve the accuracy of text segmentation.

[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Example 1

[0052] like Figure 1 As shown, this embodiment provides a scene text segmentation method, comprising the following steps:

[0053] Step 101: Use a training set to train a scene text segmentation network to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module is added pixel by pixel to the output of the coarse-grained text segmentation module as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used to perform fine-grained segmentation on the output of the coarse-grained text segmentation module.

[0054] Step 102: Use the scene text segmentation model to perform text segmentation on the image to be segmented to obtain a scene text segmentation result.

[0055] The present invention combines real scenes with text to produce a large number of synthetic data sets that meet certain requirements, including training sets. In view of the limitations of traditional convolutional neural networks for scene text segmentation, the present invention is based on the Segformer semantic segmentation method, making full use of the global context information perception ability of its Transformer encoder structure, and proposes a new scene text segmentation network (TSRNet). It also proposes multi-level feature maps (multi-level features supervision) and introduces a hybrid loss function (hybrid loss) for network supervision and training. In addition, a targeted convolution refinement module (CRM) is designed to further refine the coarse-grained segmentation results.

[0056] The overall structure of TSRNet is as follows Figure 2As shown, the SegFormer network in the coarse-grained text segmentation module is specifically the SegFormer-B1 segmentation model. The difference is that the present invention adds a multi-layer perceptron (the first multi-layer perceptron), which is used to reduce the dimension of the multi-level feature map (Multi-level features) channel output by the coarse-grained text segmentation module (Coarse Segmentation Module, CSM), reduce the parameters of the model, thereby reducing the complexity of the model and accelerating the supervised training of the model.

[0057] The SegFormer network consists of two parts: a hierarchical Transformer encoder and a lightweight multi-layer perceptron decoder. The image passes through the encoder of the SegFormer network and generates and The feature map of the original image size and the feature maps of different sizes and dimensions are passed through the decoder composed of the corresponding linear layers to generate coarse-grained text segmentation feature maps of the same size and dimension. Finally, they are spliced ​​in the dimensional direction and input into the linear layer composed of 1×1 convolution and interpolated to obtain the feature map of H×W×1. The obtained feature map is the input of the next module CRM. In order to improve the accuracy of the scene text segmentation model and make the training converge faster, 4 multi-layer perceptron (MLP) layers composed of 1×1 channel convolution are specially added. The feature maps output by the CSM module are all reduced in dimension in the channel direction to make the dimension become 1 for model training supervision. The single channel convolution operation is as follows: Figure 4 shown.

[0058] Specifically, the SegFormer network includes a first Transformer encoder, a second Transformer encoder, a third Transformer encoder, a fourth Transformer encoder, MLP Layers, and a second multi-layer perceptron connected in sequence. The outputs of the first Transformer encoder, the second Transformer encoder, and the third Transformer encoder are all connected to the input of the MLP Layers. The first Transformer encoder is used to output The second Transformer encoder is used to output the feature map of the original image size. The third Transformer encoder is used to output the feature map of the original image size. The feature map of the original size, the fourth Transformer encoder is used to output The feature map of the original image size, the original image size is the size of the image input to the first Transformer encoder. The MLP Layers is used to Feature map of the original image size, Feature map of the original image size, Feature map of original image size and The feature maps of the original image size are subjected to dimensionality reduction operations respectively, and 4 text segmentation feature maps are output.

[0059] The sample data in the training set includes input data and label data. The input data is an image combining a real scene and text, and the label data is an image after text segmentation is performed on the input data.

[0060] When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to perform supervised training on the scene text segmentation network.

[0061] The convolutional refinement module structure is similar to the U-net network, consisting of an encoder-decoder. Both the encoding and decoding paths contain a single convolution operation and four convolution blocks. The specific structure is as follows Figure 3 As shown. The convolution refinement module includes an encoder and a decoder, the output of the encoder is connected to the input of the decoder, the encoder includes a first convolution layer, a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence, the first convolution block, the second convolution block, the third convolution block and the fourth convolution all include convolution operations, downsampling operations, layer normalization operations and GELU activation function operations performed in sequence, the decoder includes a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block and a second convolution layer connected in sequence, the fifth convolution block, the sixth convolution block, the seventh convolution block and the eighth convolution block all include upsampling operations, convolution operations, layer normalization operations and GELU activation function operations performed in sequence; the first convolution layer is skip-connected to the second convolution layer, the first convolution block is skip-connected to the eighth convolution block, the second convolution block is skip-connected to the seventh convolution block, and the third convolution block is skip-connected to the sixth convolution block. The input of the first convolution layer and the output of the second convolution layer are pixel-by-pixel added as the output of the convolution refinement module.

[0062] During encoding, the output of the coarse-grained text segmentation module serves as the intermediate-layer feature map information. After being fed into the convolutional refinement module, this intermediate-layer feature map first undergoes a 3×3 convolution operation, expanding the dimension from 1D to 64D. Next, it passes through four convolutional blocks. During this process, the feature map dimension remains unchanged at 64D, without any expansion or reduction. Downsampling is achieved by changing the stride of the convolution operation. The activation function is replaced by the commonly used ReLU function with the GELU activation function, and batch normalization (BatchNorm) is replaced by layer normalization (LayerNorm). During decoding, upsampling is first performed and then fused with the feature information passed through the encoder skip link. The channel-wise concatenation is performed, doubling the original 128-dimensional dimension. Convolution is then performed again, reducing the dimension back to 64D. Finally, the initial input and final output of the CRM module are pixel-by-pixel summed and fused together to produce the TSRNet segmentation output.

[0063] The coarse-grained text segmentation module converts the input image into a segmentation probability map, while the lightweight convolutional refinement module, similar to the residual design, further refines the predicted scene text segmentation map by learning the residual between the coarse-grained segmentation probability map and the ground truth (Ground Truth, GT). The formula is as follows:

[0064] S refined =S coarse +S residual ;

[0065] Among them, S refined is the segmentation probability map output by the convolutional refinement module, S coarse The segmentation probability map output by the coarse-grained text segmentation module ( Figure 2 CoarseMap), S residual is the residual segmentation probability ( Figure 2 (RefinedMap).

[0066] The four text segmentation feature maps output by the first multi-layer perceptron are respectively recorded as the first text segmentation feature map (Sup1), the second text segmentation feature map (Sup2), the third text segmentation feature map (Sup3) and the fourth text segmentation feature map (Sup4), and the one text segmentation feature map output by the scene text segmentation network is recorded as the fifth text segmentation feature map (Sup5).

[0067] The loss function is a key component in the training of deep learning models. It measures the difference or error between the predicted output of the model and the true target value. In the field of image segmentation, the loss function measures the difference between the segmentation mask predicted by the model and its corresponding true segmentation mask. The goal of training is to minimize the loss function value, thereby producing a good segmentation model. For the text segmentation task, since there are some small-target scene text images in the dataset, there is a category imbalance phenomenon, which makes it difficult to learn the small-target text features. In order to improve the ability of TSRNet to learn small-target text features, the present invention adopts the Focal loss function as one of the loss functions; in addition, in view of the fact that scene text segmentation pays more attention to the characteristics of boundaries and their subtle structures, the present invention adopts the structural similarity index loss (StructuralSimilarityIndexLoss, SSIM) loss as one of the loss functions; considering that the intersection-over-union loss pays more attention to the foreground area (large text area) with a larger proportion, the present invention also uses it as one of the loss functions. In order to make full use of the above three loss functions, the present invention combines them to form a hybrid loss function. Focal loss is used to address the imbalance in small object text segmentation and maintain smooth gradients across all pixels. IoU is used to focus on large text regions. SSIM uses a larger loss value near boundaries to encourage predictions to maintain the structure of the original image and push background predictions toward zero, helping to focus on boundaries and foreground regions during training optimization. The hybrid loss function simultaneously supervises training for the entire segmentation network.

[0068] The hybrid loss function is expressed as:

[0069]

[0070] Among them, L represents the mixed loss, K is the number of text segmentation feature maps, K value is 5, α k The weight coefficient of the kth text segmentation feature map (the weight coefficient of the loss of the kth text segmentation feature map to the total mixed loss), l (k) is the loss value of the k-th text segmentation feature map, is the Focal loss of the k-th text segmentation feature map, is the structural similarity index loss of the k-th text segmentation feature map, is the intersection-over-union loss of the k-th text segmentation feature map.

[0071] The Focal loss function is a loss function used to solve the problem of category imbalance. The Focal loss function is expressed as:

[0072] l focal =-α t (1-p t )γ log(p t )

[0073] Among them, l focal Represents the intersection-over-union loss function value, p t represents the probability that the predicted output of the scene text segmentation network belongs to a positive sample, α t is the first weight factor, α t Specifically, it is a factor used to adjust the sample weights. γ is the second weight factor. γ is a factor used to adjust the weights of difficult and easy samples. When γ is 0, Focal loss degenerates to the standard cross-entropy loss function.

[0074] The structural similarity index loss was originally designed for image quality assessment. It captures the structural information in the image. This loss function measures the structural similarity between the predicted and true label segmentation masks, so it is often used in problems where the structural relationship between different regions in the segmentation mask is important. Therefore, it is incorporated into the training loss of this invention to learn the structural information of the true label image. First, let x = {x j :j=1,2,3,...,N 2},y={y j :j=1,2,3,...,N 2}, x, y are the pixel values ​​of the corresponding area (size N×N) intercepted from the predicted probability map (the text segmentation feature map predicted and output by the scene text segmentation model of the present invention) and the real binary segmentation mask, respectively. j Indicates the pixel value of the jth pixel on x, y j Represents the pixel value of the jth pixel on y.

[0075] The structural similarity index loss function for calculating the structural similarity index loss is expressed as:

[0076]

[0077] Among them, l ssim Represents the structural similarity index loss function value, μ x Represents the average value of the pixel value of the text segmentation feature map of the predicted output, σ x Represents the standard deviation of the pixel values ​​of the text segmentation feature map of the predicted output, μ y represents the average value of the pixel value of the real text segmentation map, σ y Represents the standard deviation of the pixel values ​​of the actual text segmentation map, σ xyRepresents the covariance (covariance of x and y) of the predicted output text segmentation feature map pixel value and the actual text segmentation map pixel value, C1 represents the first constant, C2 represents the second constant, C1 and C2 are set to any smaller constant, C1 = 0.0001, C2 = 0.0002.

[0078] In semantic segmentation, IOU loss (also known as Jaccard index or Jaccard similarity) can be used as a measure of how well the predicted segmentation map matches the true segmentation map. It is calculated as the ratio of the intersection of the predicted segmentation map and the true segmentation map to their union. Mathematically, IOU is defined as:

[0079]

[0080] Among them, the intersection is the number of pixels correctly classified as this class in the predicted and true segmentation maps, and the union is the total number of pixels belonging to this class in the predicted or true segmentation maps. Intuitively, Figure 5 shown.

[0081] IOU loss is often used in combination with other loss functions, such as cross entropy loss, to balance the trade-off between accuracy and stability during training. In this paper, IOU loss is used as one of the loss functions for training to penalize predictions that are far from the true situation and encourage the model to produce more accurate segmentation maps.

[0082] The intersection-over-union loss function is expressed as:

[0083]

[0084] Among them, l iou Represents the intersection-over-union loss function value, H represents the height of the predicted output text segmentation feature map, W represents the width of the predicted output text segmentation feature map, P(r,c) represents the pixel value at position (r,c) on the predicted output text segmentation feature map, and G(r,c) represents the pixel value at position (r,c) on the actual text segmentation map.

[0085] The CRM module in TSRNet is used only once at the original scale of the segmentation map. Furthermore, because the CSM module in TSRNet is based on the lightweight SegFormer model and incorporates a lightweight design into the implementation of the convolutional refinement module, TSRNet is more lightweight than other text segmentation networks. As shown in Table 1, the parameters and total floating-point operations of the TSRNet network model are far fewer than those of the scene text segmentation network (TexRNet) implemented based on DeeplabV3+.

[0086] Table 1 Network model parameters

[0087]

[0088] Among them, FLOPs represents the total floating point number of network calculations, 1M=1 million, 1G=1000M.

[0089] In summary, TSRNet not only outperforms other models in segmentation performance, but its coarse segmentation-refinement architecture is also simpler and easier to use, greatly saving limited hardware resources.

[0090] The implementation environment and training method process of the present invention are as follows.

[0091] The PyTorch code runs on the Linux kernel-based Ubuntu 20.04 operating system. The scene text segmentation model was trained on a computer equipped with an NVIDIA RTX 2080Ti GPU. The specific implementation environment also includes: Nvidia GeForce GTX2080Ti GPU, PyTorch version 1.2.0, TorchVision version 0.4.0, and Python version 3.7.13.

[0092] Given that the image resolution of the training dataset is relatively large at 512×512 pixels, in order to ensure good segmentation results while taking into account limited hardware resources and saving video memory, the backbone network of SegFormer adopts the SegFormer-B1 structure and uses pre-trained weights for training. The training process of the TSRNet network is divided into two stages:

[0093] In the first stage, the freezing stage, the parameters of the backbone network of SegFormer in TSRNet are frozen. At this time, the parameters of the feature extraction network do not change, the occupied video memory is small, and only the network is fine-tuned with a batch size of 4.

[0094] In the second stage, the non-frozen stage, the backbone network parameters of SegFormer are continuously updated during training, which occupies a large amount of video memory. Therefore, the batch size of training is set to half of that during frozen training, batch size = 2.

[0095] In order to more comprehensively evaluate the performance of the scene text segmentation model of the present invention, the dataset used in the present invention consists of a self-built dataset and a public dataset. The dataset pictures are as follows: Figure 6As shown. The self-built dataset refers to a self-built dataset composed of pictures generated based on the existing scene text synthesis algorithm and pictures edited using Photoshop software. The public datasets include TextSeg and MLT_S, among which TextSeg has 2646 training set pictures and 1037 test sets for training; MLT_S has a total of 5540 pictures, 3000 of which are selected as training sets and 1000 as test sets. In order to simulate the scene text pictures in real situations as much as possible, a large number of synthesized visually credible scene text images are used for the training of the current model and its downstream task models. According to the characteristics of the scene text itself, the present invention adopts data enhancement processing for the dataset pictures synthesized based on the algorithm.

[0096] The synthetic data set used in the present invention includes a total of 3060 synthetic scene text images, of which the training set consists of 2142 images, and the remaining 918 images are used as a test set. In view of the characteristics of scene text images, in order to make the synthetic data images more diverse and their distribution as close as possible to the samples in the real scene, to avoid overfitting during the model training process, thereby improving the generalization ability of the model, it is necessary to perform image enhancement processing on the synthetic scene text images, which includes some random rendering and geometric transformation of the images. Common methods include adjusting the contrast, saturation and brightness of the image and controlling the translation, rotation, size and color of the text in the image. Examples of enhanced images are as follows: Figure 7 shown.

[0097] In order to effectively evaluate the performance of the scene text segmentation model of the present invention and facilitate comparison with other segmentation network models, the present invention adopts two evaluation indicators commonly used in the segmentation field, IoU and F-Score, for evaluation and comparison.

[0098] The intersection over union (IoU) is the ratio of the intersection of the predicted segmentation map and the target area of ​​the true segmentation map to their union, which measures the degree of overlap between the predicted and true segmentation maps. The specific calculation formula is as follows:

[0099]

[0100] Among them, true positive (TP): the number of pixels correctly classified as belonging to the object of interest; false positive (FP): the number of pixels incorrectly classified as belonging to the object of interest; false negative (FN): the number of pixels that belong to the object of interest but are not correctly classified. The confusion matrix table is shown in Table 2.

[0101] Table 2 Confusion matrix

[0102]

[0103] Among them, T represents true, F represents false, P represents positive, and N represents negative.

[0104] F-Score is one of the commonly used evaluation metrics in the field of image segmentation. It combines precision and recall into one metric, providing a balance between these two important aspects of segmentation performance.

[0105] Precision measures the proportion of correctly classified pixels to the total number of pixels classified as belonging to the object of interest. It is calculated as follows:

[0106]

[0107] Recall measures the proportion of correctly classified pixels to the total number of pixels belonging to the object of interest. It is calculated as follows:

[0108]

[0109] F-Score, also known as F1 Score, is the reconciliation of precision and recall, and its calculation formula is:

[0110]

[0111] β is used to weigh the weights between precision and recall. Generally, β is set to 1, and an F-Score of 1.0 indicates perfect precision and recall, while a value of 0 indicates the worst performance. The F-Score provides a comprehensive measure of the performance of the segmentation algorithm because it considers both precision and recall in its calculation.

[0112] Comparative Experiments: To evaluate and verify the effectiveness of our scene text segmentation model, we conducted comparative experiments comparing TSRNet with DeeplabV3+, HRNetV2-W48, and a scene text segmentation network (TexRNet). Table 3 lists the performance of these networks on synthetic datasets, TextSeg, and MLT_S.

[0113] Table 3 Evaluation indicators of different segmentation networks on different datasets (IOU unit: %)

[0114]

[0115] Among them, DeeplabV3+ comes from "Encoder-decoder with atrous separableconvolution for semantic image segmentation", HRNetV2-W48 comes from "Deep high-resolution representation learning for visual recognition" and TexRNet comes from "Rethinking text segmentation: A novel dataset and a text-specific refinementapproach".

[0116] Table 4. The percentage of TSRNet that is superior to other model indicators

[0117]

[0118] As shown in Table 3, TSRNet outperforms other networks in both IOU and F-Score evaluation metrics. Table 4 shows a comparison of these metrics. Table 4 shows that TSRNet outperforms other networks in IOU and F-Score by 0.60% to 4.30% and 0.81% to 1.64% on synthetic datasets, 0.15% to 3.50% and 0.43% to 1.2% on TextSeg datasets, and 1.16% to 7.16% and 0.59% to 6.61% on MLT_S datasets. The comparative experimental results show that TSRNet has significant advantages over other segmentation networks.

[0119] Figure 8 The segmentation results of different models on the synthetic dataset, TextSeg and MLT_S are shown in the figure. The areas pointed by arrows in the second to fourth rows of the first column represent under-segmentation areas, the areas pointed by arrows in the second and third columns from the second to fourth rows represent over-segmentation areas, and the areas pointed by arrows in the fifth row of the first column and the fifth row of the third column represent normal segmentation. Figure 8 It can be seen that the TSRNet network does not have under-segmentation and over-segmentation errors compared with other networks, and the segmentation effect on the three datasets is better than other models, proving that it has obvious advantages not only in the segmentation of small target scene text but also in the segmentation of large target scene text.

[0120] To verify the effectiveness of the CRM (Convolution Refinement Module) module proposed in this paper, the multi-level features supervision, and the hybrid loss function used, the following ablation experiments were conducted. Table 5 lists the mean values ​​of the indicators of different methods on the TextSeg dataset:

[0121] Table 5 Mean indicators of different methods on the TextSeg dataset

[0122]

[0123] As shown in Table 5, in the ablation experiment, the TSRNet network achieved the best results in IoU and F-Score, which were 87.01% and 92.50% respectively. The network with only the CRM module removed achieved the second best results, which were 86.04% and 91.73% respectively. The network using only multi-level feature map supervision outperformed the basic model SegFormer, with IoU and F-Score of 85.07% and 91.51% respectively.

[0124] In order to further illustrate the contribution of CRM module, multi-level features supervision and hybrid loss function to TSRNet network. Figure 9 and Figure 10 As shown in the figure, for the basic model SegFormer, multi-level features supervision is introduced, and the IoU and F-Score increase by 0.69% and 0.26% respectively; and on this basis, the hybrid loss function is introduced, and the IoU and F-Score increase by 0.67% and 0.22% respectively; finally, the CRM module is introduced to further optimize the segmentation results, and the IoU and F-Score increase by 0.97% and 0.77% respectively, which is the largest contribution.

[0125] The ablation experiment segmentation results are shown in the figure Figure 11 As shown, Figure 11(a) shows the segmentation result of SegForm with BCE_DICE_Loss, (b) shows the segmentation result after introducing multi-level feature map supervision, (c) shows the segmentation result after introducing the hybrid loss function, and (d) shows the segmentation result after introducing the CRM module. The arrows point to areas where the segmentation results improve successively with the introduction of multi-level feature map supervision, hybrid loss function, and CRM module. The experimental results fully verify the effectiveness and necessity of the innovations proposed in this invention.

[0126] Based on the SegFormer network, the present invention makes full use of its advantages in the field of semantic segmentation and proposes a network model TSRNet for scene text segmentation. TSRNet introduces multi-level feature map supervision and hybrid loss function for network training, and introduces the CRM (Convolution Refinement Module) module to further refine and optimize the segmentation results. Comparative experiments and ablation experiments show that the TSRNet network proposed in the present invention solves the difficulties of scene text segmentation to a certain extent and achieves better segmentation results. In addition, lightweight is also its outstanding advantage. Only a computer with a 2080Ti graphics card is needed to carry out the training of the TSRNet network model.

[0127] Example 2

[0128] This embodiment provides a scene text segmentation system, including:

[0129] A scene text segmentation model training module is used to train a scene text segmentation network using a training set to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module is added pixel by pixel to the output of the coarse-grained text segmentation module as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used to perform fine-grained segmentation on the output of the coarse-grained text segmentation module.

[0130] The scene text segmentation model application module is used to use the scene text segmentation model to perform text segmentation on the image to be segmented to obtain a scene text segmentation result.

[0131] When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to perform supervised training on the scene text segmentation network.

[0132] Example 3

[0133] This embodiment provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the scene text segmentation method according to embodiment 1.

[0134] This embodiment further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the scene text segmentation method described in embodiment 1 is implemented.

[0135] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0136] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A scene text segmentation method, characterized in that: include: A scene text segmentation network is trained using a training set to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module is added pixel by pixel to the output of the coarse-grained text segmentation module as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used to perform fine-grained segmentation on the output of the coarse-grained text segmentation module; Using the scene text segmentation model to perform text segmentation on the image to be segmented, to obtain a scene text segmentation result; When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to supervise the training of the scene text segmentation network; The convolution refinement module includes an encoder and a decoder, the output of the encoder is connected to the input of the decoder, the encoder includes a first convolution layer, a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence, the first convolution block, the second convolution block, the third convolution block and the fourth convolution block all include convolution operations, downsampling operations, layer normalization operations and GELU activation function operations performed in sequence, the decoder includes a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block and a second convolution layer connected in sequence, the fifth convolution block, the sixth convolution block, the seventh convolution block and the eighth convolution block all include upsampling operations, convolution operations, layer normalization operations and GELU activation function operations performed in sequence; the first convolution layer is jump-connected to the second convolution layer, the first convolution block is jump-connected to the eighth convolution block, the second convolution block is jump-connected to the seventh convolution block, the third convolution block is jump-connected to the sixth convolution block, and the input of the first convolution layer and the output of the second convolution layer are added pixel by pixel as the output of the convolution refinement module.

2. The scene text segmentation method according to claim 1, characterized in that: The four text segmentation feature maps output by the first multi-layer perceptron are respectively recorded as the first text segmentation feature map, the second text segmentation feature map, the third text segmentation feature map and the fourth text segmentation feature map, and the one text segmentation feature map output by the scene text segmentation network is recorded as the fifth text segmentation feature map; The scene text segmentation network is trained using a mixed loss function, which is expressed as: Among them, L represents the mixed loss, K is the number of text segmentation feature maps, K value is 5, α k The weight coefficient of the kth text segmentation feature map, l (k) is the loss value of the k-th text segmentation feature map, is the Focal loss of the k-th text segmentation feature map, is the structural similarity index loss of the k-th text segmentation feature map, is the intersection-over-union loss of the k-th text segmentation feature map.

3. The scene text segmentation method according to claim 2, characterized in that: The Focal loss function for calculating the Focal loss is expressed as: l focal =-a t (1-p t ) γ log(p t ) Among them, l focal represents the intersection-over-union loss function value, p t represents the probability that the predicted output of the scene text segmentation network belongs to a positive sample, α t is the first weight factor, and γ is the second weight factor.

4. The scene text segmentation method according to claim 2, characterized in that: The structural similarity index loss function for calculating the structural similarity index loss is expressed as: in, Represents the structural similarity index loss function value, μ x Represents the average value of the pixel value of the text segmentation feature map of the predicted output, σ x Represents the standard deviation of the pixel values ​​of the text segmentation feature map of the predicted output, μ y represents the average value of the pixel value of the real text segmentation map, σ y Represents the standard deviation of the pixel values ​​of the actual text segmentation map, σ xy It represents the covariance between the pixel value of the predicted text segmentation feature map and the pixel value of the actual text segmentation map, C1 represents the first constant, and C2 represents the second constant.

5. The scene text segmentation method according to claim 2, characterized in that: The intersection-over-union loss function for calculating the intersection-over-union loss is expressed as: in, represents the intersection-over-union loss function value, H represents the height of the predicted output text segmentation feature map, W represents the width of the predicted output text segmentation feature map, P(r,c) represents the pixel value at position (r,c) on the predicted output text segmentation feature map, and G(r,c) represents the pixel value at position (r,c) on the actual text segmentation map.

6. The scene text segmentation method according to claim 1, characterized in that: The SegFormer network includes a first Transformer encoder, a second Transformer encoder, a third Transformer encoder, a fourth Transformer encoder, MLP Layers, and a second multi-layer perceptron connected in sequence, wherein the outputs of the first Transformer encoder, the second Transformer encoder, and the third Transformer encoder are all connected to the input of the MLP Layers, and the first Transformer encoder is used to output The second Transformer encoder is used to output the feature map of the original image size. The third Transformer encoder is used to output the feature map of the original image size. The feature map of the original size, the fourth Transformer encoder is used to output A feature map of the original image size, where the original image size is the size of the image input to the first Transformer encoder; The MLP Layers are used to Feature map of original image size, Feature map of original image size, The feature map of the original image size and The feature maps of the original image size are subjected to dimensionality reduction operations respectively, and 4 text segmentation feature maps are output.

7. A scene text segmentation system, characterized in that: include: A scene text segmentation model training module, used for training a scene text segmentation network with a training set to obtain a scene text segmentation model; the scene text segmentation network includes a coarse-grained text segmentation module and a convolution refinement module, the output of the coarse-grained text segmentation module is connected to the input of the convolution refinement module, and the output of the convolution refinement module is added pixel by pixel to the output of the coarse-grained text segmentation module as the output of the scene text segmentation network; the coarse-grained text segmentation module is a SegFormer network; the convolution refinement module is used for performing fine-grained segmentation on the output of the coarse-grained text segmentation module; A scene text segmentation model application module is used to use the scene text segmentation model to perform text segmentation on the image to be segmented to obtain a scene text segmentation result; When the scene text segmentation network is trained, the SegFormer network adds a first multi-layer perceptron, the output of the MLP Layers in the SegFormer network is connected to the first multi-layer perceptron, and the four text segmentation feature maps output by the first multi-layer perceptron and the one text segmentation feature map output by the scene text segmentation network are used to supervise the training of the scene text segmentation network; The convolution refinement module includes an encoder and a decoder, the output of the encoder is connected to the input of the decoder, the encoder includes a first convolution layer, a first convolution block, a second convolution block, a third convolution block and a fourth convolution block connected in sequence, the first convolution block, the second convolution block, the third convolution block and the fourth convolution block all include convolution operations, downsampling operations, layer normalization operations and GELU activation function operations performed in sequence, the decoder includes a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block and a second convolution layer connected in sequence, the fifth convolution block, the sixth convolution block, the seventh convolution block and the eighth convolution block all include upsampling operations, convolution operations, layer normalization operations and GELU activation function operations performed in sequence; the first convolution layer is jump-connected to the second convolution layer, the first convolution block is jump-connected to the eighth convolution block, the second convolution block is jump-connected to the seventh convolution block, the third convolution block is jump-connected to the sixth convolution block, and the input of the first convolution layer and the output of the second convolution layer are added pixel by pixel as the output of the convolution refinement module.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the scene text segmentation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A computer program is stored therein, and when the computer program is executed by a processor, the scene text segmentation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Semantic segmentation-based to-be-cleaned surface type identification method and cleaning equipment for executing method

    CN116152493A

  • Eye fundus image optic disk and optic cup segmentation method and device and electronic equipment

    CN116385725A