A multi-lingual handwritten word recognition and retrieval method based on visual state space network
Patent Information
- Application Number
- CN202611045615.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
现有方法通常难以同时兼顾局部笔画细节与整词长程结构,CNN类模型能够保留笔画、钩环、短横和连接点等局部证据,但全局上下文建模能力有限;Transformer类模型具有全局交互能力,但计算开销较高且常依赖大规模预训练
1. 检索精度领先:在GW、IAM、Uyghur和Esposalles四个公开基准上,完整SA-MambaSpot模型的mAP分别达到GW QBE/QBS 98.54%/99.02%、IAM QBE/QBS 93.26%/94.28%、Uyghur QBE/QBS 96.25%/95.35%、Esposalles QBE/QBS 99.72%/99.31%,在各数据集组合中均达到第一或第二水平。
Smart Images

Figure CN122821564A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, document image analysis, and handwritten word retrieval, and particularly to a multilingual handwritten word recognition and retrieval method, system, electronic device, and storage medium based on a visual state space network. Background Technology
[0002] Handwritten word retrieval aims to retrieve target word images from a collection of document images without requiring full-page transcription. It is commonly used for processing historical documents, low-resource texts, and handwritten archives. Segmented handwritten word retrieval takes pre-segmented word images as input, and the query can be a word image or a character string. Existing methods often struggle to simultaneously consider both local stroke details and the long-range structure of the entire word. CNN-type models can preserve local evidence such as strokes, hooks, short strokes, and connectors, but their global context modeling capabilities are limited. Transformer-type models have global interaction capabilities, but they have high computational costs and often rely on large-scale pre-training. Visual state-space models can model long-range dependencies with linear complexity; however, two-dimensional stroke structures may disrupt local spatial continuity when serialized along the scan path, especially in cursive, ligature, or low-resource texts such as Uyghur, where letter position and morphology, ligatures, and punctuation require high levels of local continuity. Therefore, how to simultaneously model fine-grained stroke information and the long-range structure of the entire word in a unified and efficient framework is a key technical problem in multilingual handwritten word recognition and retrieval tasks. Summary of the Invention
[0003] I. Purpose This invention proposes a stroke-aware visual state space network, SA-MambaSpot, which improves the retrieval accuracy of pre-segmented multilingual handwritten word images under both QBE and QBS protocols by combining "convolutional stem + four-directional SS2D global modeling + local stroke convolutional branch + channel-wise gated fusion + PHOC asymmetric supervision" to maintain the PHOC retrieval paradigm.
[0004] II. Technical Solution To achieve the above objectives, the present invention adopts the following technical solution: 1. Image Input and Convolution Stem The input is a pre-segmented grayscale handwritten word image, uniformly adjusted to 128×448; first, a lightweight convolution STEM is applied to preserve the stroke continuity and reduce the spatial resolution to 1 / 4 of the original, generating a visual token containing strokes, edges and character structure for subsequent long-range modeling. 2. SA-VSS backbone feature extraction The backbone uses VMamba configuration with a depth of [2, 2, 15, 2], a base width of 96, a stage channel width of [96, 192, 384, 768], and a maximum DropPath rate of 0.3. (1) The four-stage hierarchical backbone consists of stacked SA-VSS blocks. The spatial resolution is reduced and the channel width is increased through patch merging between stages, gradually expanding the receptive field and generating word-level representations for PHOC prediction. (2) Each SA-VSS block computes the global state space update G and the local convolution update L in parallel, and after gating fusion, they enter the shared residual flow and feedforward MLP sub-layer. (3) This structure injects stroke information into each state space block while keeping the PHOC retrieval process unchanged, so as to improve the stroke separability in deep features. 3. Four-way SS2D global branching The global branch in the SA-VSS block uses VMamba's two-dimensional selective scan SS2D: the feature map is expanded into a one-dimensional sequence along four complementary paths: row-first, column-first, and their reverse. After performing selective state-space recursion, it is folded back to the two-dimensional grid and fused. (1) Four-way scanning is used to capture long-range context such as character order, word length and whole word layout, and to model the whole word structure with linear complexity. (2) Since two-dimensional strokes may be separated in serialized scanning, simple SS2D may weaken the evidence of local continuous strokes. Therefore, this invention adds explicit local stroke branches in its bypass. 4. Local stroke convolution branch (1) The local stroke branch uses lightweight depthwise separable convolution to model the local stroke neighborhood in each channel. Specifically, it consists of 3×3 depthwise convolution, batch normalization, GELU nonlinearity, Dropout, and 1×1 pointwise projection. This branch is not a preprocessing step before SS2D scanning, but an independent path parallel to the global branch, used to extract fine-grained stroke evidence such as hooks, loops, short horizontal strokes, and connection points. (2) This branch only adds a small number of parameters and computational cost. In the experiment, the complete model only added 3.61 M parameters and 1.44 G FLOPs to the baseline. 5. Adaptive Parallel Fusion Instead of setting residuals for each branch separately, the fused features are entered into the shared residual connection, reducing competition between branches and maintaining the stability of the backbone feature evolution. (1) Given a global update G and a local update L, channel-wise gating F=G+σ(g)⊙L is used for fusion, where g is the learnable channel gating and ⊙ represents channel-wise broadcast multiplication. (2) The gating is initialized to α0=0.1, so that each block is initially approximately the standard VSS block. During the training process, the intensity of local stroke information injected into different channels is adaptively learned. 6. PHOC Embedding and Retrieval The PHOC head performs 2D LayerNorm, global average pooling, and linear classification on the features of the last layer of the backbone, and outputs a PHOC attribute probability vector via Sigmoid, which is used as a word-image embedding. In instance-based retrieval (QBE), the query vector is the PHOC probability vector predicted from the query image; in string-based retrieval (QBS), the query vector is the standard binary PHOC vector of the query string. Image library samples are sorted in ascending order by cosine distance. 7. Asymmetric PHOC monitoring and system implementation (1) During the training phase, asymmetric loss (ASL) is used to supervise sparse PHOC multi-label attributes. γ−=4, γ+=0, and pruning threshold m=0.05 are set to reduce the influence of negatively resistant samples and enhance the differentiation of confusing attributes. (2) A multilingual handwritten word recognition and retrieval system, including an image preprocessing module, an SA-VSS feature extraction module, a PHOC attribute embedding module, an ASL training module, a cosine distance ranking module and a result output module; (3) An electronic device and a computer-readable storage medium, wherein the processor executes a program to implement the above method steps.
[0005] III. Beneficial Effects 1. Leading Retrieval Accuracy: On four public benchmarks—GW, IAM, Uyghur, and Esposalles—the full SA-MambaSpot model achieved mAP of 98.54% / 99.02% for GW QBE / QBS, 93.26% / 94.28% for IAM QBE / QBS, 96.25% / 95.35% for Uyghur QBE / QBS, and 99.72% / 99.31% for Esposalles QBE / QBS, respectively, ranking first or second in each dataset combination. 2. Local and global synergy: The SA-VSS block combines the whole-word context modeling capability of four-directional SS2D with local stroke convolution branches. Ablation experiments show that on IAM, it improves from 91.25% / 93.55% of the CNN-VMamba-PHOC baseline to 93.26% / 94.28% of the full model. 3. Low static complexity: The full model has 53.44 M parameters, 19.87 G FLOPs, and 268.56 MiB peak memory, which is lower than the external baseline in terms of parameter count, FLOPs, and peak memory; local stroke branches only bring about 7% to 8% additional overhead. 4. Effective supervision of sparse attributes: PHOC prediction is a sparse multi-label problem. ASL improves the QBE / QBS mAP in IAM ablation from 92.44% / 93.71% to 93.26% / 94.28% by reducing the contribution of negatively oriented samples and retaining positive attribute learning. 5. Wide range of applicable texts: Experiments cover scenarios such as handwritten English and Uyghur, and it is suitable for instance-based retrieval and string-based retrieval of pre-segmented handwritten word images, and can be further extended to more low-resource texts. Attached Figure Description
[0006] Figure 1 This is a diagram of the overall network architecture of the SA-MambaSpot of this invention. Figure 2 SA-VSS module structure diagram Figure 3 A schematic diagram of the Top-5 search results for SA-MambaSpot on the IAM and Uyghur datasets. Figure 4 A schematic diagram of the module ablation results of SA-MambaSpot on the IAM dataset. Figure 5 This diagram illustrates the efficiency comparison results of SA-MambaSpot on the IAM dataset. Detailed Implementation Plan The feasibility of the present invention is further illustrated below with reference to embodiments and experimental data, but this does not constitute a limitation of the present invention.
[0007] Example 1 1. Dataset: IAM English handwriting dataset, containing over 115,000 word instances from 657 writers. It is divided according to the official settings, using 6,161 lines for training, 1,840 lines for validation, and 1,861 lines for testing, and adopts a leave-one-out setting during retrieval. 2. Training configuration: Input grayscale word image size is 128×448; optimizer is AdamW; training is performed on an NVIDIA GeForce RTX 3090 Ti for 80,000 iterations, batch size is 64, and weight decay is 1.5×10⁻⁶. -4 No learning rate warmup; the learning rate is 1×10 for the first 40,000 iterations. -4 Up to 60,000 iterations, it becomes 5 × 10 -5 After that, it becomes 1×10 -5 Use 36 Latin symbols from levels {2, 3, 4, 5}, and add 50 of the most frequent bigrams from level 2. 3. Results: The complete SA-MambaSpot achieved QBE mAP=93.26% and QBS mAP=94.28% on IAM. The complete model had 53.44 M parameters, 19.87 G FLOPs, and a peak memory usage of 268.56 MiB. 4. Ablation Experiments: Module ablation experiments were conducted on the IAM dataset. The QBE / QBS mAP(%) results are as follows: VMamba (w / o CNN stem, w / o PHOC): 85.63% / — CNN-VMamba (w / o PHOC): 86.75% / — VMamba-PHOC (w / o CNN stem): 90.11% / 94.26% CNN-VMamba-PHOC (base): 91.25% / 93.55% +Local branching, additive fusion, BCE: 92.01% / 93.42% +Local branching, gated fusion of this invention, BCE: 92.44% / 93.71% Full SA-MambaSpot: 93.26% / 94.28% It is evident that local stroke modeling, parallel gating fusion, and ASL supervision together contribute to the cumulative gain.
[0008] Example 2 1. Datasets: Four publicly available handwritten word retrieval benchmarks: GW, IAM, Uyghur, and Esposalles, covering scenarios such as English handwriting, historical documents, and low-resource cursive text. 2. Training configuration: Input grayscale word image size is 128×448; optimizer is AdamW; weight decay is 1.5×10⁻⁶. -4 No learning rate warmup; all datasets use the PHOC attribute embedding paradigm; the Latin dataset uses 36 Latin symbols at levels {2,3,4,5}, and adds 50 of the most frequent bigrams at level 2; the Uyghur dataset uses 37 Uyghur symbols at levels {1,2,4,8}, and does not use bigram attributes due to their cursive and position-dependent glyph characteristics. 3. Comparative Experiment: (1) On the GW dataset, the SA-MambaSpot full model achieved QBE mAP=98.54% and QBS mAP=99.02%. (2) On the IAM dataset, the SA-MambaSpot full model achieved QBE mAP=93.26% and QBS mAP=94.28%, with the QBE index slightly higher than that of HWNet v2, which was pre-trained using large-scale synthetic data. (3) On the Uyghur dataset, the SA-MambaSpot full model achieved QBE mAP=96.25% and QBS mAP=95.35%, verifying its applicability to low-resource cursive text. (4) On the Esposalles dataset, the SA-MambaSpot full model achieved QBE mAP=99.72% and QBS mAP=99.31%, indicating that the method remains robust in historical document retrieval scenarios.
[0009] in conclusion This invention achieves collaborative modeling of local stroke details and long-range structure of whole words in pre-segmented handwritten word image scenarios through SA-VSS stroke-aware visual state space module, local stroke convolutional branch, channel-wise gated fusion, and ASL-supervised PHOC head. Experiments show that the method consistently outperforms the CNN-VMamba-PHOC baseline on four benchmarks: GW, IAM, Uyghur, and Esposalles, and achieves retrieval performance comparable to or better than existing representative methods. It can be used for instance-based and string-based retrieval tasks of word images in multilingual handwritten documents.
Claims
1. A multilingual handwritten word recognition and retrieval method based on a visual state space network, characterized in that, Includes the following steps: S1: Obtain a pre-segmented multilingual handwritten word image dataset, and perform grayscale conversion, size normalization, and retrieval sample organization processing on the handwritten word images; S2: Construct the stroke-aware visual state space network SA-MambaSpot, which includes a convolutional stem, a four-stage hierarchical backbone network formed by stacking multiple SA-VSS blocks, and a PHOC attribute prediction head. S3: Input the handwritten word image into the convolutional stem to preserve stroke continuity and reduce spatial resolution, and then input it into the SA-VSS block. The long-range structure of the whole word is modeled through four-directional SS2D state space branches, and local stroke neighborhood features are extracted through parallel local stroke convolutional branches. S4: The local stroke neighborhood features are adaptively injected into the global state space representation using a learnable channel-by-channel gating mechanism, and the fused features are fed into a shared residual stream and a feedforward network to obtain a word-level visual representation. S5: The word-level visual representation is mapped to a PHOC attribute probability vector through the PHOC attribute prediction head, and the PHOC attribute prediction head is trained using asymmetric loss. S6: During the retrieval phase, the query image or query string is converted into a query PHOC vector, the image library handwritten word image is converted into an image library PHOC embedding, the cosine distance between the query PHOC vector and the image library PHOC embedding is calculated, and the Top-K retrieval results are returned in ascending order of distance.
2. The method according to claim 1, wherein the SA-VSS block comprises: (1) Global SS2D state space branch, used to perform two-dimensional selective scanning of the input feature map and obtain global state space update; (2) Local stroke convolution branch, used to extract stroke detail features in the local neighborhood through lightweight depthwise separable convolution; (3) An adaptive fusion unit is used to fuse the global state space update and the local stroke detail features according to the learnable channel-by-channel gating weights.
3. The method according to claim 2, characterized in that, The global SS2D state space branch expands the input feature map into four one-dimensional sequences along the row-first direction, the column-first direction, and the corresponding two reverse directions. After performing selective state space recursion, these sequences are folded back into a two-dimensional grid and fused to model the long-range dependencies of character order, word length, and whole word layout.
4. The method according to claim 1, wherein the local stroke convolution branch comprises: (1) Two-dimensional normalization unit, used to normalize the input features; (2) Depth convolution unit, used to aggregate local stroke neighborhood information in each channel by k×k depth convolution; (3) Nonlinear transformation and projection unit, used to generate local convolution updates that match the global state space update through batch normalization, GELU activation, random deactivation and 1×1 pointwise convolution.
5. The method according to claim 4, characterized in that, The local stroke convolution branch is an independent path parallel to the SS2D state space branch, rather than a preprocessing path before SS2D scanning. It is used to inject local stroke evidence into the global representation after the state space scan.
6. The method according to claim 1, characterized in that, The asymmetric loss is used for PHOC sparse multi-label attribute supervision, and its expression is: L ASL =(-1 / D)∑[y d (1-p d ) γ+ log(p d )+(1-y d )p m,d γ- log(1-p m,d )] Where p m,d =max(p d -m, 0), where D is the dimension of the PHOC attribute, p d Let y be the predicted probability of the d-th PHOC attribute. d γ is the corresponding true attribute label, γ+ and γ- are the asymmetric focusing coefficients, and m is the negative sample pruning threshold. During training, γ+=0, γ-=4, and m=0.05 are set to reduce the contribution of easily negative samples and enhance the distinction of confusing PHOC attributes.
7. A system based on the method of any one of claims 1 to 6, characterized in that, include: (1) Image processing module, used to execute S1; (2) Visual state space feature extraction module, used to execute S2 and S3; (3) Dual-branch feature adaptive fusion module, used to execute S4; (4) PHOC attribute training module, used to execute S5; (5) Handwritten word retrieval module, used to execute S6 and output Top-K retrieval results.
8. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the method as claimed in any one of claims 1 to 6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 6.