Chinese handwritten form detection method and system based on text kernel and text clustering

Through feature extraction and dynamic feature fusion based on ResNet-18, combined with text clustering module, the accuracy and robustness of multi-scale text detection in the prior art are solved, and efficient detection of multi-scale, dense arrangement and arbitrary shape text is achieved.

CN120279569APending Publication Date: 2025-07-08HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357666.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing text detection methods are sensitive to multi-scale texts, and there is a problem that the loss of small text details and misjudgment of large text structures coexist. In dense text instances are prone to boundary adhesions, existing pixel clustering methods are difficult to ensure instance distinction, and fixed weight feature fusion strategies cannot adapt to the dynamic changes in text spatial distribution.

Method used

The Chinese handwritten text feature map is extracted based on the ResNet-18 feature extraction module, the number of channels is adjusted through the convolution layer, and the multi-scale text feature enhancement module and dynamic feature fusion module are combined. The spatial attention mechanism is used to dynamically fuse the feature maps of different scales, and the similarity vectors between text pixels are learned through the text clustering module, the text pixels are aggregated on the corresponding text core, and the text area, text kernel and text cluster loss function are introduced for model training.

Benefits of technology

It improves the accuracy and robustness of text detection, can effectively handle multi-scale, dense arrangement and arbitrary shape text in complex scenarios, significantly improving the detection accuracy of small text details and large text structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279569A_ABST
    Figure CN120279569A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese handwritten form detection method and system based on a text kernel and text clustering. The method comprises the following steps: acquiring a Chinese handwritten form text image; extracting Chinese handwritten text feature maps, and adjusting the number of channels to unify the number of channels of each feature map; the deep layer features are gradually restored to a higher resolution ratio through up-sampling operation, and the deep layer features are fused with shallow layer features of corresponding scales; dynamically fusing feature maps of different scales; learning similarity vectors among text pixels, and aggregating the text pixels to corresponding text kernels; using the predicted similarity vector to guide the aggregation of text pixels; and adding the text region loss, the text kernel loss and the text clustering loss into model training, and carrying out iteration by utilizing a gradient descent algorithm until the maximum number of iterations is reached or the model is converged. According to the technical scheme, the text kernel prediction technology and the dynamic pixel clustering technology are fused, and the text detection precision and robustness are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to a Chinese handwritten text detection method and system based on text kernels and text clustering. Background Art

[0002] Current text detection methods are sensitive to multi-scale texts, and there are problems of coexistence of loss of small text details and misjudgment of large text structures. And for dense text instances, boundary adhesion is likely to occur, and it is difficult for existing pixel clustering methods to ensure the case of instance distinguishability. In addition, the fixed-weight feature fusion strategy cannot adapt to the dynamic changes in the spatial distribution of texts. Summary of the Invention

[0003] The main object of the present invention is to propose a Chinese handwritten text detection method and system based on text kernels and text clustering, aiming to improve the text detection accuracy and robustness for multi-scale, densely arranged and arbitrarily shaped texts in complex scenarios by fusing text kernel prediction and dynamic pixel clustering technologies.

[0004] To achieve the above object, the present invention provides a Chinese handwritten text detection method based on text kernels and text clustering, and the method includes the following steps:

[0005] Step S10, obtaining a Chinese handwritten text image,

[0006] Step S20, extracting a Chinese handwritten text feature map based on a ResNet-18 feature extraction module, adjusting the number of channels through a convolutional layer so that the number of channels of each feature map is unified, and mapping feature maps of different scales to a unified feature space:

[0007] Step S30, on the basis of feature mapping, gradually restoring deep features to a higher resolution through an upsampling operation based on a multi-size text feature enhancement module, and fusing them with shallow features of corresponding scales;

[0008] Step S40, dynamically fusing feature maps of different scales based on a dynamic feature fusion module by using a spatial attention mechanism;

[0009] Step S50, learning a similarity vector between text pixels based on a text clustering module, and aggregating text pixels onto corresponding text kernels, thereby realizing the detection of arbitrarily shaped text instances;

[0010] Step S60, using the predicted similarity vector based on the text clustering module to guide the aggregation of text pixels;

[0011] Step S70: Incorporate the text region loss, text kernel loss, and text clustering loss into the model training, and use the gradient descent algorithm for iteration until the maximum number of iterations is reached or the model converges.

[0012] A further technical solution of the present invention is that the step S20 includes:

[0013] Extract Chinese handwritten text feature maps based on the ResNet-18 feature extraction module, and obtain features {c2, c3, c4, c5}. The number of channels of each feature map is adjusted through convolutional layers conv2, conv3, conv4, conv5 respectively, so that the number of channels of each feature map is unified to C. This process is expressed as:

[0014] conv2 = Conv2d(c2, C) (1-1);

[0015] conv3 = Conv2d(c3, C) (1-2);

[0016] conv4 = Conv2d(c4, C) (1-3);

[0017] conv5 = Conv2d(c5, C) (1-4);

[0018] Where Conv2d represents the convolution operation, and C is the target number of channels.

[0019] A further technical solution of the present invention is that the step S30 is implemented by layer-by-layer upsampling and element-wise addition. The specific steps are as follows:

[0020] out4 = up5(conv5) + conv4 (1-5);

[0021] out3 = up4(out4) + conv3 (1-6);

[0022] out2 = up3(out3) + conv2 (1-7);

[0023] Where up5, up4, up3 are upsampling modules that restore the resolution of the feature maps by using the nearest neighbor interpolation or bilinear interpolation method.

[0024] A further technical solution of the present invention is that the multi-scale text enhancement module introduces a feature refinement module at each scale, further extracts features through convolution operations, and restores the feature maps to a higher resolution through upsampling operations. The feature refinement module is expressed as:

[0025] p5 = refine5(conv5) (1-8);

[0026] p4 = refine4(out4) (1-9);

[0027] p3 = refine3(out3) (1-10);

[0028] p2 = refine2(out2) (1-11);

[0029] Among them, refine5, refine4, refine3, and refine2 are feature refinement modules, which are composed of convolutional layers and upsampling operations.

[0030] A further technical solution of the present invention is that, on the basis of feature refinement, the multi-scale text feature enhancement module splices feature maps of different scales through splicing in the channel dimension to obtain a comprehensive multi-scale feature map. This process is expressed as:

[0031] fuse = cat(p5, p4, p3, p2)(1-12);

[0032] Among them, cat represents the splicing operation in the channel dimension.

[0033] A further technical solution of the present invention is that the step S40 includes:

[0034] Assume that the input feature map is composed of N feature maps of different scales, denoted as First, adjust these feature maps of different scales to the same resolution, then splice them, and input the spliced feature map into a 3×3 convolutional layer to obtain intermediate features This process can be expressed by the formula:

[0035] S = Conv(concat([X , X1,..., X N-1 )) (1-13);

[0036] Among them, Conv represents a 3×3 convolutional operation, and concat represents the splicing operation of feature maps;

[0037] Next, use the spatial attention module to process the intermediate feature S to calculate the attention weights The implementation of the spatial attention module includes two parts: attention calculation in the spatial dimension and attention calculation in the feature dimension; first, compress the feature map S into a single-channel global feature map through global average pooling, then extract the attention weights in the spatial dimension through a convolutional network, and finally fuse the spatial attention weights with the input feature map by element-wise addition; this process is expressed by the formula:

[0038] global x = mean(S, dim = 1, keepdim = True) (1-14);

[0039] global x = Spatial_Wise(global x ) + S (1-15);

[0040] A = Attention_Wise(global x ) (1-16);

[0041] Among them, Spatial_Wise is a network containing two convolutional layers, used to extract the attention weights in the spatial dimension; Attention_Wise is a convolutional layer, used to fuse the spatial attention weights with the input feature map and generate the final attention weight A:

[0042] Finally, the attention weight A is split into N parts according to the channel dimension, and each part is weighted and multiplied with the corresponding feature map X i to obtain the fused feature map The formula is:

[0043]

[0044] Among them, A i represents the attention weight at the i-th scale, and ⊙ represents element-wise multiplication.

[0045] A further technical solution of the present invention is that the step S50 includes:

[0046] For each text pixel p, the network predicts its similarity vector F(p), indicating the possibility that the pixel belongs to a certain text instance. Among them, the calculation formula of the similarity vector is:

[0047] F(p) = W p · f(p) (1-18);

[0048] Among them, f(p) is the feature vector of pixel p, and W p is the learned weight matrix;

[0049] For each text kernel k i , its similarity vector G(k i ) is calculated by the average value of the similarity vectors of all pixels within the kernel:

[0050]

[0051] Among them, K i is the one belonging to kernel ki The set of all text pixels, |k i | is the number of elements in the set;

[0052] To ensure that text pixels can be correctly aggregated onto the corresponding text kernels, the PA module introduces an aggregation loss function L agg , and this loss function L agg is achieved by minimizing the distance between text pixels and the corresponding kernels:

[0053]

[0054] where N is the number of text instances, T i is the set of all pixels of the i-th text instance, and D(p, k i ) is the distance between text pixel p and kernel k i , which is defined as:

[0055] D(p, k i ) = max(||F(p) - G(k i )|| - δ agg , 0) 2 (1 - 21);

[0056] where δ agg is a constant used to aggregate text kernel pixels into the inner text kernel;

[0057] To maintain the distinctiveness between different text kernels, the text clustering module introduces a discriminative loss function L dis to ensure that the distance between different text kernels is large enough:

[0058]

[0059] where D(k i , k j ) is the distance between kernel k i and kernel k j , which is defined as:

[0060] D(k i , k j ) = max(δ dis - ||G(k i ) - G(k j )||, 0) 2 (1 - 23);

[0061] where δ dis is a constant used to control the minimum distance between different kernels;

[0062] The total loss function L of the text clustering module cIt is the aggregation loss function L agg and the discriminant loss function L dis The sum of:

[0063] L c = L agg + L dis (1-24).

[0064] A further technical solution of the present invention is that the step S60 includes:

[0065] Nuclear detection: Find the connected regions in the nuclear segmentation result, and each connected region is regarded as a text nucleus;

[0066] Pixel aggregation: For each text nucleus k i , merge the adjacent text pixels p onto this nucleus, provided that the Euclidean distance between their similarity vectors is less than the threshold d;

[0067] Iterative aggregation: Repeat the above steps until there are no eligible adjacent text pixels.

[0068] A further technical solution of the present invention is that the step S70 includes:

[0069] Express the overall loss function as:

[0070] L = L text + αL kernel + βL c (1-25);

[0071] Among them, L text is the text region loss function, L kernel is the text nucleus loss function, L c is the text clustering loss function, and α and β are hyperparameters used to balance the weights of each part of the loss;

[0072] The text region loss function L text is used to supervise the segmentation result of the text region, and this loss function adopts the Dice loss formula:

[0073]

[0074] Among them, P text (i) and G text (i) respectively represent the values of the predicted text region and the true text region at the i-th pixel point;

[0075] The text nucleus loss function L kernel is used to supervise the segmentation result of the text nucleus, and also adopts the Dice loss formula:

[0076]

[0077] Among them, P kernel (i) and G kernel (i) respectively represent the values of the predicted text kernel and the true text kernel at the i-th pixel. The true text kernel is obtained by shrinking the original polygon annotation.

[0078] To achieve the above object, the present invention also proposes a Chinese handwritten character detection system based on text kernel and text clustering. The system includes a memory, a processor, and a Chinese handwritten character detection program based on text kernel and text clustering stored on the processor. When the Chinese handwritten character detection program based on text kernel and text clustering is run by the processor, it executes the steps of the method as described above.

[0079] The Chinese handwritten character detection method and system based on text kernel and text clustering of the present invention effectively improve the text detection accuracy and robustness by the above technical solutions, integrating text kernel prediction and dynamic pixel clustering technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.

[0081] Figure 1 is a schematic flowchart of a preferred embodiment of the Chinese handwritten character detection method based on text kernel and text clustering of the present invention;

[0082] Figure 2 is a schematic framework diagram involved in the Chinese handwritten character detection method based on text kernel and text clustering of the present invention;

[0083] Figure 3 is a schematic diagram of the working principle of the dynamic feature fusion module.

[0084] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0086] For the detection scenario of multi-scale, densely arranged, and arbitrarily shaped text in complex scenarios, the present invention proposes a Chinese handwritten text detection method based on text kernels and text clustering. The main technical solution adopted is to improve the text detection accuracy and robustness by integrating text kernel prediction and dynamic pixel clustering techniques. The execution entities involved in the Chinese handwritten text detection method based on text kernels and text clustering of the present invention include a multi-scale text feature enhancement module, a dynamic feature fusion module, and a text clustering module.

[0087] Please refer to Figure 1 and Figure 2 , the preferred embodiment of the Chinese handwritten text detection method based on text kernels and text clustering of the present invention includes the following steps:

[0088] Step S10, obtain a Chinese handwritten text image.

[0089] Step S20, extract the Chinese handwritten text feature map based on the ResNet-18 feature extraction module, and adjust the number of channels through the convolutional layer so that the number of channels of each feature map is unified, and map the feature maps of different scales to a unified feature space.

[0090] Step S30, on the basis of feature mapping, based on the multi-size text feature enhancement module, gradually restore the deep features to a higher resolution through upsampling operations and fuse them with the shallow features of the corresponding scales.

[0091] Step S40, based on the dynamic feature fusion module, dynamically fuse the feature maps of different scales using the spatial attention mechanism.

[0092] Step S50, based on the text clustering module, learn the similarity vectors between text pixels, aggregate the text pixels onto the corresponding text kernels, so as to realize the detection of arbitrarily shaped text instances.

[0093] Step S60, based on the text clustering module, use the predicted similarity vectors to guide the aggregation of text pixels.

[0094] Step S70, add the text region loss, text kernel loss, and text clustering loss to the model training, and use the gradient descent algorithm for iteration until the maximum number of iterations is reached or the model converges.

[0095] Specifically, the step S20 includes:

[0096] Extract the Chinese handwritten text feature map based on the ResNet-18 feature extraction module, and obtain the features {c2, c3, c4, c5}, respectively adjust the number of channels through the convolutional layers (conv2, conv3, conv4, conv5) so that the number of channels of each feature map is unified to C. This process is expressed as:

[0097] conv2 = Conv2d(c2, C) (1-1);

[0098] conv3 = Conv2d(c3, C) (1-2);

[0099] conv4 = Conv2d(c4, C) (1-3);

[0100] conv5 = Conv2d(c5, C) (1-4).

[0101] Among them, Conv2d represents the convolution operation, and C is the number of target channels. Through this process, feature maps of different scales are mapped to a unified feature space, providing a basis for subsequent fusion operations.

[0102] Based on the feature mapping, the multi-scale text feature enhancement module gradually restores the deep features to a higher resolution through upsampling operations and fuses them with the shallow features of the corresponding scale. This process is achieved through layer-by-layer upsampling and element-wise addition, and the specific steps are as follows:

[0103] out4 = up5(conv5) + conv4 (1-5);

[0104] out3 = up4(out4) + conv3 (1-6);

[0105] out2 = up3(out3) + conv2 (1-7);

[0106] Among them, up5, up4, up3 are upsampling modules, and usually the resolution of the feature map is restored by means of nearest neighbor interpolation (nearest) or bilinear interpolation (bilinear). Through layer-by-layer upsampling and fusion, the model can effectively combine feature information of different scales and enhance the perception ability of multi-scale targets.

[0107] To further optimize the feature representation, the multi-scale text feature enhancement module introduces feature refinement modules at each scale. These modules further extract features through convolution operations and restore the feature maps to a higher resolution through upsampling operations. Specifically, the feature refinement modules can be expressed as:

[0108] p5 = refine5(conv5) (1-8);

[0109] p4 = refine4(out4) (1-9);

[0110] p3 = refine3(out3) (1-10);

[0111] p2 = refine2(out2) (1-11);

[0112] Among them, refine5, refine4, refine3, and refine2 are feature refinement modules, usually composed of convolutional layers and upsampling operations. These modules not only further extract features but also restore the feature maps to a higher resolution through upsampling operations for multi-scale splicing.

[0113] Based on feature refinement, the multi-scale text feature enhancement module splices feature maps of different scales through splicing in the channel dimension to obtain a comprehensive multi-scale feature map. This process can be expressed as:

[0114] fuse = cat(p5, p4, p3, p2) (1-12);

[0115] Among them, cat represents the splicing operation in the channel dimension. Through this process, the model integrates feature information of different scales into one feature map.

[0116] In semantic segmentation and scene text detection tasks, features of different scales play different roles in describing text instances. Shallow or large-scale features can capture the details of small text instances but are difficult to grasp the global information of large text instances; while deep or small-scale features can better perceive the overall structure of large text instances but may lose the details of small text instances. To give full play to the advantages of features of different scales, this paper proposes a dynamic feature fusion module, and its working principle is as Figure 3 shown.

[0117] Specifically, in this embodiment, the step S40 includes:

[0118] Assume that the input feature map is composed of N feature maps of different scales, denoted as (In practical applications, N is usually set to 4). First, adjust these feature maps of different scales to the same resolution, and then splice them. The spliced feature map is input into a 3×3 convolutional layer to obtain intermediate features This process can be expressed by the formula:

[0119] S = Conv(concat([X o , X1,..., X N-1 )) (1-13);

[0120] Among them, Conv represents the 3×3 convolutional operation, and concat represents the splicing operation of feature maps.

[0121] Next, the intermediate feature S is processed using the spatial attention module to calculate the attention weights. The implementation of the spatial attention module includes two main parts: attention calculation in the spatial dimension and attention calculation in the feature dimension. Specifically, first, the feature map S is compressed into a single-channel global feature map through global average pooling, and then a convolutional network is used to extract the attention weights in the spatial dimension. Finally, the spatial attention weights are fused with the input feature map by element-wise addition. This process can be expressed by the formula:

[0122] global x = mean(S, dim = 1, keepdim = True) (1-14);

[0123] global x = Spatial_Wise(global x ) + S (1-15);

[0124] A = Attention_Wise(global x ) (1-16);

[0125] Among them, Spatial_Wise is a network containing two convolutional layers for extracting the attention weights in the spatial dimension; Attention_Wise is a convolutional layer for fusing the spatial attention weights with the input feature map and generating the final attention weight A.

[0126] Finally, the attention weight A is split into N parts according to the channel dimension, and each part is weighted and multiplied with the corresponding feature map, resulting in the fused feature map. The formula is:

[0127]

[0128] Among them, A i represents the attention weight at the i-th scale, and ⊙ represents element-wise multiplication.

[0129] Through the above operations, the dynamic feature fusion module can dynamically fuse feature maps of different scales. The spatial attention mechanism plays a key role in this process, making the attention weights more flexible in the spatial dimension, enabling the model to adaptively adjust the attention weights according to the spatial distribution of text instances. This fusion method not only retains the detailed information of small text instances but also enhances the overall perception ability of large text instances, significantly improving the performance of the model in multi-scale text detection tasks.

[0130] In the handwritten text detection task, a text clustering module is proposed to handle the aggregation of text pixels for accurately reconstructing complete text instances. The text clustering module aggregates text pixels onto corresponding text kernels by learning the similarity vectors between text pixels, thereby achieving the detection of arbitrarily shaped text instances.

[0131] Specifically, in this embodiment, the step S50 includes:

[0132] Calculation of similarity vectors:

[0133] For each text pixel p, the network predicts its similarity vector F(p), which represents the likelihood that the pixel belongs to a certain text instance. The calculation formula for the similarity vector is:

[0134] F(p) = W p ·f(p) (1 - 18);

[0135] where f(p) is the feature vector of pixel p, and W p is the learned weight matrix.

[0136] For each text kernel k i , its similarity vector G(k i ) is calculated as the average of the similarity vectors of all pixels within the kernel:

[0137]

[0138] where K i is the set of all text pixels belonging to kernel k i , and |k i | is the number of elements in the set.

[0139] To ensure that text pixels can be correctly aggregated onto the corresponding text kernels, the PA module introduces an aggregation loss function L agg . This loss function is achieved by minimizing the distance between text pixels and the corresponding kernels:

[0140]

[0141] where N is the number of text instances, T i is the set of all pixels of the i-th text instance, D(p, k i ) is the distance between text pixel p and kernel k i , and is defined as:

[0142] D(p, k i ) = max(||F(p) - G(k i )|| - δ agg , 0) 2(1-21);

[0143] Among them, δ agg is a constant used to aggregate text kernel pixels into the text kernel, and its value is generally 0.5.

[0144] To maintain the distinctiveness between different text kernels, the text clustering module also introduces a discriminative loss function L dis to ensure that the distance between different text kernels is large enough:

[0145]

[0146] Among them, D(k i , k j ) is the distance between kernel k i and kernel k j , defined as:

[0147] D(k i , k j ) = max(λ dis - ||G(k i ) - G(k j )||, 0) 2 (1-23);

[0148] Among them, δ dis is a constant used to control the minimum distance between different kernels, which is mainly set to 3 in this article.

[0149] The total loss function L c of the text clustering module is the sum of the aggregation loss function L agg and the discriminative loss function L dis :

[0150] L c = L agg + L dis (1-24).

[0151] In the test phase, the text clustering module uses the predicted similarity vector to guide the aggregation of text pixels. The specific steps of step S60 include:

[0152] Kernel detection: Find the connected regions in the kernel segmentation result, and each connected region is regarded as a text kernel;

[0153] Pixel aggregation: For each text kernel k i , merge the adjacent text pixels p onto this kernel on the condition that the Euclidean distance between their similarity vectors is less than the threshold d;

[0154] Iterative aggregation: Repeat the above steps until there are no eligible adjacent text pixels.

[0155] Through the above method, the text clustering module can effectively handle the problem of text pixel aggregation in handwritten recognition, and improve the model's detection ability for text instances of arbitrary shapes.

[0156] Post-processing

[0157] To generate the prediction result, first calculate the image connectivity graph according to the kernel map, and generate the text region using a breadth-first search-like method based on the similarity matrix.

[0158] In this embodiment, the loss function consists of three parts, the text region loss L text , the text kernel loss L kernel , and the text clustering loss L c .

[0159] The specific steps of step S70 include:

[0160] Express the overall loss function as:

[0161] L = L text + αL kernel + βL c (1 - 25);

[0162] Among them, L text is the text region loss function, L kernel is the text kernel loss function, L c is the text clustering loss function, and α and β are hyperparameters used to balance the weights of each part of the loss. In the experiment, α is set to 0.5 and β is set to 0.25.

[0163] The text region loss function L text is used to supervise the segmentation result of the text region, mainly aiming at the problem of the imbalance in the number between text pixels and non-text pixels. This loss function adopts the Dice loss formula:

[0164]

[0165] Among them, P text (i) and G text (i) respectively represent the values of the predicted text region and the real text region at the i-th pixel point. The real text region is a binary image, where text pixels are 1 and non-text pixels are 0.

[0166] The text kernel loss function L kernel is used to supervise the segmentation result of the text kernel, and also adopts the Dice loss formula:

[0167]

[0168] Among them, P kernel(i) and G kernel (i) represents the values of the predicted text kernel and the ground-truth text kernel at the i-th pixel respectively, and the ground-truth text kernel is obtained by shrinking the original polygon annotation.

[0169] In the specific implementation process of the Chinese handwritten text detection method based on text kernel and text clustering in this embodiment, for the data preprocessing step, a Resnet-18 pre-trained model is used for feature extraction. The experimental results (Accuracy, Recall, F1-score) on the table detection dataset xfund[1] are shown in Table 1.

[0170] Table 1

[0171] Acc(%) Recall(%) F1-score(%) PANNET[2] 75.3 79.5 77.3 DBNET[3] 76.4 80.2 78.3 Ours 80.1 82.4 81.2

[0172] [1]Xu Y, Lv T, Cui L, et al. XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding[J]. Findings of the Association for Computational Linguistics: ACL 2022, 2022: 3214 - 3224.

[0173] In summary, the Chinese handwritten text detection method based on text kernel and text clustering of the present invention, through the above technical solutions, combines text kernel prediction and dynamic pixel clustering technology, effectively improving the text detection accuracy and robustness.

[0174] To achieve the above object, the present invention also proposes a Chinese handwritten text detection system based on text kernel and text clustering. The system includes a memory, a processor, and a Chinese handwritten text detection program based on text kernel and text clustering stored on the processor. When the Chinese handwritten text detection program based on text kernel and text clustering is run by the processor, it executes the steps of the method described in the above embodiment, which will not be elaborated here.

[0175] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the concept of the present invention using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of the present invention.

Claims

1. A Chinese handwritten character detection method based on a text kernel and text clustering, characterized in that, The method includes the following steps: Step S10: Obtain a Chinese handwritten text image; Step S20: Extract the Chinese handwritten text feature map based on the ResNet-18 feature extraction module, adjust the number of channels through the convolutional layer to make the number of channels of each feature map unified, and map feature maps of different scales to a unified feature space; Step S30: On the basis of feature mapping, based on the multi-scale text feature enhancement module, gradually restore the deep features to a higher resolution through upsampling operations and fuse them with the shallow features of the corresponding scale; Step S40: Based on the dynamic feature fusion module, adopt a spatial attention mechanism to dynamically fuse feature maps of different scales; Step S50: Based on the text clustering module, learn the similarity vector between text pixels, aggregate the text pixels onto the corresponding text kernels, so as to realize the detection of text instances of arbitrary shapes; Step S60: Based on the text clustering module, use the predicted similarity vector to guide the aggregation of text pixels; Step S70: Add the text region loss, text kernel loss, and text clustering loss to the model training, and use the gradient descent algorithm for iteration until the maximum number of iterations is reached or the model converges.

2. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 1, wherein The said step S20 includes: Extract the Chinese handwritten text feature map based on the ResNet-18 feature extraction module, and obtain features {c2, c3, c4, c5}. Respectively adjust the number of channels through the convolutional layers conv2, conv3, conv4, convS to make the number of channels of each feature map unified to C. This process is expressed as: conv2 = Conv2d(c2, C) (1-1); conv3 = Conv2d(c3, C) (1-2); conv4 = Conv2d(c4, C) (1-3); conv5 = Conv2d(c5, C) (1-4); Wherein, Conv2d represents the convolution operation, and C is the target number of channels.

3. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 2, wherein The said step S30 is realized through layer-by-layer upsampling and element-wise addition. The specific steps are: out4 = up5(conv5)+conv4 (1-5); out3 = up4(out4)+conv3 (1-6); out2 = up3(out3)+conv2 (1-7); Wherein, up5, up4, up3 are upsampling modules that restore the resolution of the feature map by using the nearest neighbor interpolation or bilinear interpolation method.

4. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 3, wherein The multi-scale text enhancement module introduces a feature refinement module at each scale, further extracts features through convolution operations, and restores the feature map to a higher resolution through upsampling operations. The feature refinement module is expressed as: p5 = refine5(conv5) (1-8); p4 = refine4(out4) (1-9); p3 = refine3(out3) (1-10); p2 = refine2(out2) (1-11); Among them, refine5, refine4, refine3, and refine2 are feature refinement modules, which are composed of convolutional layers and upsampling operations.

5. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 4, wherein Based on feature refinement, the multi-scale text feature enhancement module stitches feature maps of different scales through stitching in the channel dimension to obtain a comprehensive multi-scale feature map. This process is expressed as: fuse = cat(p5, p4, p3, p2) (1-12); Among them, cat represents the stitching operation in the channel dimension.

6. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 5, wherein Step S40 includes: Assume that the input feature maps consist of N feature maps of different scales, denoted as First, adjust these feature maps of different scales to the same resolution, and then splice them. The spliced feature maps are input into a 3×3 convolutional layer to obtain intermediate features This process can be expressed by the formula: S = Conv(concat([X0, X1,..., X N-1 )) (1-13); Among them, Conv represents a 3×3 convolutional operation, and concat represents the stitching operation of feature maps; Next, the intermediate feature S is processed using the spatial attention module to calculate the attention weights. The implementation of the spatial attention module includes two parts: attention calculation in the spatial dimension and attention calculation in the feature dimension. First, the feature map S is compressed into a single-channel global feature map through global average pooling, then a convolutional network is used to extract the attention weights in the spatial dimension, and finally, the spatial attention weights are fused with the input feature map by element-wise addition. This process is expressed by the formula: global x = mean(S, dim=1, keepdim=True) (1 - 14); global x = Spatial_Wise(global x ) + S(1 - 15); A = Attention_Wise(global x ) (1 - 16); Among them, Spatial_Wise is a network containing two convolutional layers for extracting spatial dimension attention weights; Attention_Wise is a convolutional layer for fusing spatial attention weights with the input feature map and generating the final attention weight A; Finally, split the attention weight A into N parts according to the channel dimension, and perform weighted multiplication on each part with the corresponding feature map, to obtain the fused feature map The formula is as follows: Among them, A i represents the attention weight of the i-th scale, and ⊙ represents element-wise multiplication.

7. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 6, wherein Step S50 includes: For each text pixel p, the network predicts its similarity vector F(p), indicating the possibility that the pixel belongs to a certain text instance. Among them, the calculation formula of the similarity vector is: F(p) = W p ·f(p) (1-18); where f(p) is the feature vector of pixel p, and W p is the learned weight matrix; For each text kernel k i , its similarity vector G(k i ) is calculated as the average of the similarity vectors of all pixels within the kernel: where K i is the set of all text pixels belonging to the kernel k i , |k i | is the number of elements in the set; To ensure that text pixels can be correctly aggregated onto the corresponding text kernels, the PA module introduces an aggregation loss function L agg , and this loss function L agg is achieved by minimizing the distance between text pixels and the corresponding kernels: where N is the number of text instances, T i is the set of all pixels of the i-th text instance, D(p, k i ) is the distance between the text pixel p and the kernel k i is defined as: D(p, k i ) = max(||F(p) - G(k i )|| - δ agg , 0) 2 (1 - 21); where δ agg is a constant for aggregating text kernel pixels into the text kernel; To maintain the distinctiveness between different text kernels, the text clustering module introduces a discriminative loss function L dis , to ensure that the distances between different text kernels are large enough: Among them, D(k i , k j ) is the distance between kernel k i and kernel k j , and is defined as: D(k i , k j ) = max(δ dis - ||G(k i ) - G(k j )||, 0) 2 (1 - 23); where δ dis is a constant used to control the minimum distance between different cores; The total loss function L of the text clustering module c is the aggregation loss function L agg and the discriminative loss function L dis and is: L c = L agg + L dis (1 - 24).

8. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 7, wherein, Step S60 includes: Nucleus detection: Find the connected regions in the nucleus segmentation result, and each connected region is regarded as a text nucleus; Pixel aggregation: For each text kernel k i , neighboring text pixels p are merged onto the kernel if the Euclidean distance between their similarity vectors is less than a threshold d; Iterative aggregation: Repeat the above steps until there are no neighboring text pixels that meet the conditions.

9. The Chinese handwritten character detection method based on a text kernel and text clustering according to claim 8, characterized in that, Step S70 includes: Express the overall loss function as: L = L text + αL kernel + βL c (1 - 25); Among them, L text is the text area loss function, L kernel is the text kernel loss function, L c is the text clustering loss function, and α and β are hyperparameters used to balance the weights of the losses of each part; Text region loss function L text To supervise the segmentation result of the text region, this loss function adopts the Dice loss formula: Among them, P text (i) and G text (i) respectively represent the values of the predicted text region and the ground-truth text region at the i-th pixel point; Text kernel loss function L kernel To supervise the segmentation result of the text kernel, the Dice loss formula is also used: where P kernel (i) and G kernel (i) represent the values of the predicted text kernel and the ground-truth text kernel at the i-th pixel, respectively, and the ground-truth text kernel is obtained by shrinking the original polygon annotation.

10. A Chinese handwritten character detection system based on a text kernel and text clustering, characterized in that, The system includes a memory, a processor, and a Chinese handwritten character detection program based on text kernels and text clustering stored on the processor. When the Chinese handwritten character detection program based on text kernels and text clustering is run by the processor, it executes the steps of the method according to any one of claims 1 to 9.