Text Detection Method Based on Word Focus Network and Sample Weight Allocation

By using the method of independent allocation of word focus backbone network and sample weights in scene text detection, the problem of slow model training speed and poor extreme aspect ratio text detection is solved, and fast training and high-performance detection are achieved.

CN116844146BActive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310805134.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-03
Publication Date
2025-06-13
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

In the prior art, the model training speed is slow and the extreme aspect ratio is poorer than text examples.

Method used

The scene text detection method based on word focus backbone network and independent allocation of sample weights is adopted. By integrating the spatial information of text instances into model training, the model convergence speed is accelerated, and the channel attention and spatial attention mechanism are used to help the model pay attention to pixel points containing semantic information. Finally, the label allocation strategy of independently allocating weights is used to help the model learn difficult samples.

Benefits of technology

Fast model training and efficient text detection are achieved, especially when handling extreme aspect ratio text instances, the performance is significantly improved, and the model robustness is also improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844146B_ABST
    Figure CN116844146B_ABST
Patent Text Reader

Abstract

In view of the problems such as slow model training speed and poor detection effect of extreme aspect ratio text instances by existing methods, the present invention proposes a scene text detection method based on a word-focus backbone network and independent allocation of sample weights, integrates the spatial information of text instances into the early model training to accelerate the model convergence speed, and uses channel attention and spatial attention mechanisms to help the model focus on pixel points containing semantic information. Finally, a label assignment strategy with independently allocated weights is used to help the model learn difficult samples to improve the robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision understanding, and particularly relates to a scene text detection method based on a word-focus backbone network and independent allocation of sample weights. Background Art

[0002] In recent years, artificial intelligence technology has developed rapidly. Using deep learning to process some natural scene texts in life, that is, natural scene text detection and recognition has become a popular technology. Natural scene text detection and recognition is a very important research field in the fields of computer vision and artificial intelligence. It mainly studies whether a machine can correctly understand a picture, so as to complete the detection and recognition of targets in the picture. Summary of the Invention

[0003] Aiming at problems such as slow model training speed and poor detection effect of existing methods for text instances with extreme aspect ratios, the present invention proposes a scene text detection method based on a word-focus backbone network and independent allocation of sample weights, integrates the spatial information of text instances into the early model training, accelerates the model convergence speed, and uses channel attention and spatial attention mechanisms to help the model focus on pixel points containing semantic information. Finally, a label assignment strategy with independently allocated weights is used to help the model learn difficult samples to improve the robustness of the model.

[0004] The present invention can realize the detection and recognition of natural scene texts using deep learning, and this method has a fast inference speed and better performance than other existing methods.

[0005] First, it uses a natural scene text image dataset, inputs the pictures in the dataset into the word-focus backbone network, integrates the spatial information of natural images into the backbone network, accelerates model convergence, adds deformable convolutions in cross-layer connections, so that the network can better handle the change of feature map scale; then the obtained feature map is input into the adaptive feature screening network, and a method of fusing spatial attention and channel attention is used to help the model adaptively screen features containing semantic information of text instances, and a residual structure is used to fuse it with the original features; then the fused features are input into the probability map generation head, the probability map generation head is used to predict the pixel points of text instances, and the pixel aggregation algorithm is used to obtain the boundaries of text instances; finally, a label assignment strategy of strengthening consistency is used to help the model learn difficult samples and improve the robustness of the model.

[0006] The technical solution specifically adopted by the present invention to solve its technical problems is:

[0007] A scene text detection method based on a word-focus backbone network and independent allocation of sample weights, comprising the following steps;

[0008] Step S1: Obtain a natural scene text image dataset. Input the images in the dataset into the word-focus backbone network, integrate the spatial information of the natural images into the backbone network to accelerate model convergence, and add deformable convolutions in the cross-layer connections so that the network can better handle the changes in the scale of the feature maps;

[0009] Step S2: Input the obtained feature maps into the adaptive feature screening network. Use the method of fusing spatial attention and channel attention to help the model adaptively screen the features containing the semantic information of text instances, and use the residual structure to fuse the screened features with the original features;

[0010] Step S3: Input the fused features into the probability map generation head. Use the probability map generation head to predict the pixel points of text instances, and use the pixel aggregation algorithm to obtain the boundaries of text instances.

[0011] Further, step S1 specifically includes the following steps;

[0012] Step S11: Obtain a publicly available natural scene text dataset;

[0013] Step S12: Also record the corresponding text in the text area of the dataset into a json file for convenient subsequent recognition;

[0014] Step S13: Input the images into the word-focus backbone network in batches; obtain four feature maps of different scales, and their sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively;

[0015] The word-focus backbone network is based on the ResNet50 structure. In the Stage0 layer, the spatial information of the image is input into the structure composed of a convolutional layer - batch normalization layer - GeLU activation function - max pooling layer together with the original input, and then the input of Stage0 is input into Stage1, Stage2, Stage3, and Stage4; replace the Bottleneck structure in Stage2, Stage3, and Stage4, that is, the structure composed of 1×1 convolution, 3×3 convolution, and 1×1 convolution, with the structure of 1×1 convolution, 3×3 deformable convolution, and 1×1 convolution. Figure 1

[0016] Further, step S2 specifically includes the following steps;

[0017] Step S21: Uniformly scale the feature maps output in step S13 to 1 / 4 of the size of the original image and input them into the adaptive feature screening network; the adaptive feature screening network is composed of a convolution layer, a spatial attention module, and a channel attention module, as shown in the following formula:

[0018] F Attention_out =Conv3×3 (ChannelAttention(SpaceAttention(F neck_in ))) Formula 1

[0019] Among them, F neck_in is the input feature of the adaptive feature screening network, Conv 3×3 is a convolutional layer with a convolutional kernel size of 3×3, F Attention_out is the feature obtained after the input feature passes through the channel attention module and the spatial attention module. ChannelAttention is the channel attention module, and SpaceAttention is the spatial attention module; the channel attention module performs spatial average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature to obtain the attention weight in the channel dimension; the spatial attention module performs channel average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature to obtain the attention weight in the spatial dimension;

[0020] Step S22: Add the adaptively screened feature F Attention_out to the input feature F neck_in and input it into a 3×3 convolutional layer to obtain the fused feature F neck_out , as shown in Formula 2:

[0021] F neck_out = Conv 3×3 (F neck_in + F Attention_out ) Formula 2.

[0022] Furthermore, step S3 specifically includes the following steps;

[0023] Step S31: Use the differentiable binarization function to predict the pixels of the text instance; as shown in Formula 3:

[0024]

[0025] Among them, B i,j represents the threshold map generated by the differentiable binarization function, T i,j represents the dynamic threshold map learned by the network, P i,j represents the probability map generated by the model, and i, j represent the horizontal and vertical coordinates of the corresponding map, which are used to fit the approximate binarization function to the standard binarization curve; to distinguish the text instance from the background and use the pixel aggregation algorithm to predict the boundary of the text instance;

[0026] Step S32: Train the model using the loss function, and the loss function is as shown in Formula 4:

[0027] L = L cls + 5L T Formula Four

[0028] Among them, L cls is composed of cross - entropy loss, and the threshold map loss L T is composed of the L1 loss between the text instance boundary and the predicted pixels. The L1 loss is shown in Formula Five:

[0029]

[0030] Among them, R D represents the text boundary interval calculated by the Vatti clipping algorithm, y i represents the coordinate on the interval, x i represents the pixel coordinate predicted by the model; and R D is obtained by subtracting the polygon G obtained by shrinking the annotation information through the Vatti clipping algorithm from the polygon G obtained by dilating the annotation information through the Vatti clipping algorithm, as shown in Formula Six: D minus the polygon G obtained by shrinking the annotation information through the Vatti clipping algorithm S as shown in Formula Six:

[0031] R D = G D - G S Formula Six

[0032] The offset D for shrinking or dilating the polygon by the Vatti clipping algorithm is defined as Formula Seven:

[0033]

[0034] Among them, A is the area of the polygon, r is the change ratio, and L is the perimeter of the polygon.

[0035] Furthermore, in step S3, a label assignment strategy with enhanced consistency is used to help the model learn difficult samples to improve the robustness of the model.

[0036] Furthermore, the label assignment strategy with enhanced consistency specifically includes the following steps;

[0037] To assign greater weights to pixels with higher classification scores and position consistency; the consistency measure t of the pixel pos is defined as Formula Eight:

[0038] t pos = score × L1 α Formula Eight

[0039] Among them, score is the classification score of the pixel, L1 is the L1 loss between the predicted pixel and the text instance, and α is a hyperparameter; to enhance the variance of the positive weights, an exponential modulation factor is added, and the positive weight wpos Defined as Formula Nine:

[0040]

[0041] Where β is a hyperparameter; for negative weights, the inconsistency measure t neg Defined as Formula Ten:

[0042] t neg =-L1 γ Formula Ten

[0043] Where γ is a hyperparameter, and the negative weight w neg Defined as Formula Eleven:

[0044] w neg =t neg ×score δ Formula Eleven

[0045] Where δ is a hyperparameter, and finally w pos and w neg are substituted into the cross-entropy formula L cls , as shown in Formula Twelve:

[0046] L cls =-w pos ×ln(score)-w neg ×ln(1 - score) Formula Twelve.

[0047] Compared with the prior art, the present invention and its preferred embodiments have the following beneficial effects:

[0048] 1. The proposed scene text detection method based on a word-focus backbone network and independent assignment of sample weights has a faster inference speed and better performance compared to other existing methods.

[0049] 2. The word-focus backbone network in the present invention can accelerate the training speed of the model, and it can converge after only 100 training rounds.

[0050] 3. The adaptive feature screening network in the present invention uses a method that combines channel attention and spatial attention, which can relate the local features and global features of text instances in the image, and solve the problem of low detection rate of texts with extremely high aspect ratios by ordinary methods.

[0051] 4. The performance of the detection model can be further optimized by methods such as data augmentation, data enhancement, and model integration, and the accuracy can be further improved. Brief Description of the Drawings

[0052] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0053] Figure 1 It is a schematic diagram of the design route and principle of an embodiment of the present invention. Specific embodiments

[0054] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below and described in detail as follows:

[0055] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0056] As Figure 1 shown, this embodiment provides a scene text detection method based on a word-focus backbone network and independent assignment of sample weights, which specifically includes the following steps;

[0057] Step S1: Obtain a natural scene text image dataset, input the pictures in the dataset into the word-focus backbone network, integrate the spatial information of the natural image into the backbone network to accelerate model convergence, and add deformable convolutions in the cross-layer connection so that the network can better handle the change of the feature map scale;

[0058] Step S2: Input the obtained feature map into the adaptive feature screening network, use the method of fusing spatial attention and channel attention to help the model adaptively screen the features containing the semantic information of text instances, and use the residual structure to fuse them with the original features;

[0059] Step S3: Input the fused features into the probability map generation head, use the probability map generation head to predict the pixel points of text instances, and use the pixel aggregation algorithm to obtain the boundaries of text instances;

[0060] Step S4: Use the label assignment strategy of enhancing consistency to help the model learn difficult samples and improve the robustness of the model.

[0061] Among them, step S1 specifically includes the following steps;

[0062] Step S11: Obtain a publicly available natural scene text dataset, such as ICDAR2013, ICDAR2015, ICDAR2019, CTW1500, etc.;

[0063] Step S12: Record the corresponding text in the text area of the dataset into a json file for convenient subsequent recognition. The format of the json file is {'xxx.jpg': {'points': [[coordinates of text area 1], [coordinates of text area 2], …]}, …}.

[0064] Step S13: Input the images into the word-focus backbone network batch by batch. The word-focus backbone network is similar to the ResNet50 structure. At the Stage0 layer, in this embodiment, the spatial information of the image and the original Figure 1 are input into the structure consisting of a convolutional layer - batch normalization layer - GeLU activation function - max pooling layer, and then the input of Stage0 is input into Stage1, Stage2, Stage3, and Stage4. In this embodiment, the Bottleneck structure in Stage2, Stage3, and Stage4, that is, the structure consisting of 1×1 convolution, 3×3 convolution, and 1×1 convolution, is replaced with the structure of 1×1 convolution, 3×3 deformable convolution, and 1×1 convolution. Finally, four feature maps with different scale sizes are obtained, and their sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively.

[0065] Step S2 specifically includes the following steps;

[0066] Step S21: Uniformly scale the feature maps output in Step S13 to 1 / 4 of the size of the original image and input them into the adaptive feature screening network. The adaptive feature screening network consists of a convolutional layer, a spatial attention module, and a channel attention module, and the specific principle is shown in Formula 1:

[0067] F Attention_out =Conv 3×3 (ChannelAttention(SpaceAttention(F neck_in ))) # Formula 1

[0068] where F neck_in is the input feature of the adaptive feature screening network, Conv 3×3 is a convolutional layer with a convolutional kernel size of 3×3, F Attention_out is the feature obtained after the input feature passes through the channel attention module and the spatial attention module, ChannelAttention is the channel attention module, and SpaceAttention is the spatial attention module. The channel attention module performs spatial average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature, and finally the attention weight in the channel dimension is obtained. Similarly, the spatial attention module performs channel average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature, and finally the attention weight in the spatial dimension is obtained.

[0069] Step S22: Add the adaptively screened feature F Attention_out to the input feature F neck_in and input it into a 3×3 convolutional layer to obtain the fused feature F neck_out , as shown in Formula 2:

[0070] F neck_out = Conv 3×3 (F neck_in + F Attention_out ) # Formula 2;

[0071] Step S3 specifically includes the following steps;

[0072] Step S31: Use a differentiable binarization function to predict the pixels of the text instance. As shown in Formula 3:

[0073]

[0074] Where B i,j represents the threshold map generated by the differentiable binarization function, T i,j represents the dynamic threshold map learned by the network, P i,j represents the probability map generated by the model, and i, j represent the horizontal and vertical coordinates of the corresponding map. It is used to better fit the standard binarization curve with the approximate binarization function. Through the differentiable binarization function, the model can better distinguish the text instance from the background, and can also achieve a better segmentation effect for text instances with a relatively dense distribution. Finally, the boundary of the text instance is predicted using the pixel aggregation algorithm.

[0075] Step S32: Train the model using the loss function. The loss function is as shown in Formula 4:

[0076] L = L cls + 5L T # Formula 4;

[0077] Where, L cls is composed of the cross-entropy loss, and the threshold map loss L T is composed of the L1 loss between the text instance boundary and the predicted pixels. The L1 loss is as shown in Formula 5:

[0078]

[0079] Where R D represents the text boundary interval calculated by the Vatti clipping algorithm, y i represents the coordinate on the interval, x i represents the pixel coordinate predicted by the model. And R D is obtained by subtracting the polygon G D obtained by shrinking the annotation information through the Vatti clipping algorithm from the polygon G S obtained by dilating the annotation information through the Vatti clipping algorithm, as shown in Formula 6:

[0080] R D = G D-G S # Formula VI;

[0081] For the Vatti clipping algorithm, the offset D for shrinking or expanding the polygon is defined as Formula VII:

[0082]

[0083] where A is the area of the polygon, r is the change ratio, usually set to 0.4, and L is the perimeter of the polygon.

[0084] Step S4 specifically includes the following steps;

[0085] Step S41: Use a label strategy with enhanced consistency to help the model learn difficult samples. Specifically, it is to assign a larger weight to the pixels with higher classification scores and position consistency. In this embodiment, the consistency measure t of the pixels pos is defined as Formula VIII:

[0086] t pos = score × L1 α # Formula VIII;

[0087] where score is the classification score of the pixel, L1 is the L1 loss between the predicted pixel and the text instance, and α is a hyperparameter. To enhance the variance of the positive weights, an exponential modulation factor is added, and the positive weight w pos is defined as Formula IX:

[0088]

[0089] where β is a hyperparameter. For the negative weights, in this embodiment, the inconsistency measure t neg is defined as Formula X:

[0090] t neg = -L1 γ # Formula X;

[0091] where γ is a hyperparameter. In this embodiment, the negative weight w neg is defined as Formula XI:

[0092] w neg = t neg × score δ # Formula XI;

[0093] where δ is a hyperparameter. Finally, w pos and w neg are substituted into the cross-entropy formula L cls , as shown in Formula XII:

[0094] L cls = -w pos×ln(score) - w neg ×ln(1 - score) # Formula XII;

[0095] Specifically, in view of the problems of slow model training speed and poor detection effect of existing methods for text instances with extreme aspect ratios, the present invention proposes a scene text detection method based on a word-focus backbone network and independent assignment of sample weights, integrates the spatial information of text instances into the early model training to accelerate the model convergence speed, and uses channel attention and spatial attention mechanisms to help the model focus on pixel points containing semantic information. Finally, a label assignment strategy with independently assigned weights is used to help the model learn difficult samples and improve the robustness of the model.

[0096] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0097] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0098] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the process Figure 1 in one process or a plurality of processes and / or boxes Figure 1 steps for the functions specified in one box or a plurality of boxes.

[0100] This patent is not limited to the above best implementation. Under the inspiration of this patent, anyone can derive various other forms of scene text detection methods based on word-focus backbone networks and independent assignment of sample weights. All equivalent changes and modifications made in accordance with the scope of the patent application of the present invention shall fall within the scope covered by this patent.

Claims

1. A text detection method based on a word focus network and sample weight assignment, characterized in that, it includes the following steps; Step S1: Obtain a natural scene text image dataset, input the pictures in the dataset into the word focus backbone network, integrate the spatial information of the natural image into the backbone network to accelerate model convergence, and add deformable convolutions in the cross-layer connections so that the network can better handle changes in the feature map scale; Step S2: Input the obtained feature map into the adaptive feature screening network, and use the method of fusing spatial attention and channel attention to help the model adaptively screen the features containing the semantic information of text instances, and use the residual structure to fuse the screened features with the original features; Step S3: Input the fused features into the probability map generation head, use the probability map generation head to predict the pixel points of text instances, and use the pixel aggregation algorithm to obtain the boundaries of text instances: including step S31: Use a differentiable binarization function to predict the pixels of text instances; as shown in Equation 3: Among them, B i,j represents the threshold map generated by using the differentiable binarization function, and T i,j represents the dynamic threshold map learned by the network, and P i,j represents the probability map generated by the model. i and j represent the horizontal and vertical coordinates of the corresponding map, which are used to fit the approximate binarization function to the standard binarization curve; in order to separate text instances from the background, the pixel aggregation algorithm is used to predict the boundaries of text instances.

2. The text detection method based on a word focus network and sample weight assignment according to claim 1, characterized in that: Step S1 specifically includes the following steps; Step S11: Obtain a public natural scene text dataset; Step S12: Record the corresponding text in the text area of the dataset into a json file for convenient subsequent recognition; Step S13: Input the images into the word focus backbone network in batches; obtain four feature maps of different scales, and their sizes are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively; The word focus backbone network is based on the ResNet50 structure. At the Stage0 layer, the spatial information of the image and the original image are input into a structure composed of a convolutional-batch normalization layer-GeLU activation function-maximum pooling layer, and then the input of Stage0 is input into Stage1, Stage2, Stage3, and Stage4; Replace the Bottleneck structure in Stage2, Stage3, and Stage4, that is, the structure composed of 1×1 convolution, 3×3 convolution, and 1×1 convolution, with the structure of 1×1 convolution, 3×3 deformable convolution, and 1×1 convolution.

3. The text detection method based on a word focus network and sample weight assignment according to claim 2, characterized in that: Step S2 specifically includes the following steps; Step S21: Uniformly scale the feature map output in Step S13 to 1 / 4 of the original image size and input it into the adaptive feature screening network; the adaptive feature screening network is composed of a convolution, a spatial attention module, and a channel attention module, as shown in the following formula: F Attention_out = Conv 3×3 (ChannelAttention(SpaceAttention(F neck_in ))) Equation 1 Among them, F neck_in is the input feature of the adaptive feature screening network, Conv 3×3 is a convolutional layer with a convolutional kernel size of 3×3, and F Attention_out is the feature obtained after the input feature passes through the channel attention module and the spatial attention module. ChannelAttention is the channel attention module, and SpaceAttention is the spatial attention module; the channel attention module performs spatial average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature to obtain the attention weight in the channel dimension; the spatial attention module performs channel average pooling on the feature map, and after passing through the convolutional layer and the GeLU activation function layer, it is added to the original input feature to obtain the attention weight in the spatial dimension; Step S22: Add the adaptively filtered feature F Attention_out to the input feature F neck_in and input it into a 3×3 convolutional layer to obtain the fused feature F neck_out , as shown in Equation 2: F neck_out = Conv 3×3 (F neck_in + F Attention_out ) Equation 2.

4. The text detection method based on a word focus network and sample weight assignment according to claim 3, characterized in that: Step S3 specifically further includes the following steps; Step S32: Train the model using a loss function, and the loss function is as shown in Equation 4: L′ = L cls +5L T Formula Four Among them, L cls is composed of cross-entropy loss, and the threshold map loss L T is composed of the L1 loss between the text instance boundary and the predicted pixels. The L1 loss is shown in Equation (5): Where R D represents the text boundary interval calculated by the Vatti clipping algorithm, y i represents the coordinate on the interval, x i represents the pixel coordinate predicted by the model; and R D is obtained by subtracting the polygon G D obtained by shrinking the annotation information through the Vatti clipping algorithm from the polygon G S obtained by inflating the annotation information through the Vatti clipping algorithm, as shown in Formula 6: R D = G D - G S Formula Six The offset D for shrinking or expanding the polygon by the Vatti clipping algorithm is defined as Equation 7: where A is the area of the polygon, r is the change ratio, and L is the perimeter of the polygon.

5. The text detection method based on a word focus network and sample weight assignment according to claim 4, characterized in that: In step S3, a label assignment strategy for enhancing consistency is used to help the model learn difficult samples so as to improve the robustness of the model.

6. The text detection method based on a word focus network and sample weight assignment according to claim 5, characterized in that: The label assignment strategy for enhancing consistency specifically includes the following steps; To assign greater weights to pixels with higher classification scores and position consistency; the consistency measure t of a pixel pos is defined as Equation (8): t pos = score × L1 α Formula VIII where score is the classification score of the pixel, L1 is the L1 loss between the predicted pixel and the text instance, and α is a hyperparameter; to enhance the variance of the positive weights, an exponential modulation factor is added, and the positive weight w pos is defined as Equation (9): where β is a hyperparameter; For negative weights, the inconsistency measure t neg is defined as Equation Ten: t neg = -L1 γ Formula Ten where γ is a hyperparameter, and the negative weight w neg is defined as Equation (11): w neg = t neg × score δ Formula XI where δ is a hyperparameter, and finally substitute w pos and w neg into the cross-entropy formula L cls , as shown in Equation (12): L cls = -w pos × ln(score) - w neg × ln(1 - score) Formula XII.

Citation Information

Patent Citations

  • Natural scene text detection method and system, storage medium and computing equipment

    CN116229445A

  • Transmission tower signboard text detection and identification method based on STN-pan network

    CN116343188A