Scene text recognition method based on visual Mamba and Transformer hybrid architecture

By adopting a hybrid architecture method of visual Mamba and Transformer in scene text recognition, combining multi-branch convolutional structure and attention mechanism, the problem of low accuracy of text recognition in complex scenes is solved, and higher recognition accuracy and robustness are achieved.

CN119964176BActive Publication Date: 2025-06-06NORTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510449785.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-06
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

When the prior art faces complex scenes (such as blur, bending, occlusion, etc.), the accuracy of scene text recognition is low, making it difficult to effectively deal with complex and changeable scene text.

Method used

The scene text recognition method based on the hybrid architecture of visual Mamba and Transformer is adopted. Through the combination of spatial conversion network STN, ResNet, feature enhancement module, sequence modeling module and decoder, the feature and global context information of different scales are captured to improve the recognition accuracy.

Benefits of technology

It effectively improves the accuracy and robustness of the text recognition algorithm, can better process complex scene text, and enhances the ability to model text images in sequence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964176B_ABST
    Figure CN119964176B_ABST
Patent Text Reader

Abstract

The present application relates to a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer, including: obtaining a text image dataset; constructing a scene text recognition network; training the scene text recognition network based on the text image dataset to obtain a trained scene text recognition network; inputting the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, selecting the text corresponding to the maximum value of the corresponding probability of all text categories as the recognition result of the character. The present application uses visual Mamba to effectively compress and model the visual context, and is successfully used in fields such as visual prediction; combined with the multi-head attention mechanism of Transformer, it improves the ability to perceive global context information, enhances the ability to model the sequence of text images, and effectively improves the accuracy of the text recognition algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition, and in particular, to a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer. Background Art

[0002] Scene text refers to text content that appears in various natural environments, such as billboards, product labels, road signs, printed clothing, banners, etc., which can reflect information about a specific environment or entity. Scene text recognition is of great significance in obtaining text information from natural scenes and is one of the research hotspots in the field of computer vision. In recent years, with the rapid development of deep learning, many scene text recognition algorithms have emerged, and have achieved remarkable results in text detection and recognition tasks, and have been successfully applied to many fields such as office automation, image search, instant translation, and robot navigation.

[0003] Before the era of deep learning, text recognition algorithms mainly used segment-based recognition methods or decomposed the recognition process into multiple sub-problems. These methods mainly rely on manually designed low-level or mid-level image features, which require tedious and repetitive pre-processing and post-processing steps. Due to the limitations of manual feature representation capabilities and the complexity of the processing flow, these methods perform poorly when dealing with complex scenes, such as the blurred images in the ICDAR 2015 dataset. With the diversification of text scenes, traditional scene text recognition technology faces challenges in dealing with complex and changing scenes. Many scenes have complex text backgrounds, and the text and background are not distinguishable. In addition, the text may be in artistic fonts or handwritten fonts with different shapes, which greatly increases the difficulty of text analysis and processing. With the improvement of computer computing power, the research and application of deep learning related theories have made significant progress and achieved excellent results in various fields. Text recognition algorithms based on deep learning have gradually become mainstream. Text recognition technology based on deep learning does not require manual design of features, can handle more complex text scenes, and has higher accuracy and robustness. Although many scene text recognition algorithms have achieved outstanding results, their recognition accuracy is still low when faced with complex scenes (such as blur, curvature, occlusion, etc.). Summary of the invention

[0004] In order to overcome at least one deficiency in the prior art, the present application provides a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer.

[0005] In the first aspect, a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer is provided, comprising:

[0006] Obtain a text image dataset; samples in the text image dataset include scene text images in a variety of different fonts, different backgrounds, and different noise conditions;

[0007] Build a scene text recognition network, which includes a spatial transformer network STN, ResNet, a feature enhancement module, a sequence modeling module, and a decoder;

[0008] The scene text recognition network is trained based on the text image dataset to obtain a trained scene text recognition network; the training process includes:

[0009] The sample is input into the scene text recognition network, and the spatial transformer network STN performs geometric transformation correction on the sample to obtain the corrected image; ResNet extracts features from the corrected image to obtain local visual features; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features; the sequence modeling module uses visual Mamba and Transformer to capture the long-range dependencies and global context information in the scene text to obtain processed features for the enhanced features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probability of all text categories is selected as the recognition result of the character;

[0010] The scene text image to be recognized is input into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character.

[0011] In one embodiment, the feature enhancement module includes a first convolution branch, a second convolution branch, a third convolution branch, and an attention mechanism;

[0012] The first convolution branch includes two convolution layers connected in sequence, and the second convolution branch has the same structure as the third convolution branch, including three convolution layers connected in sequence and a dilated convolution layer;

[0013] The first convolution branch, the second convolution branch, and the third convolution branch output features of different scales. Features of different scales are input into the convolution layer. The convolution results are then input into the attention mechanism for weighted processing to obtain enhanced features.

[0014] In one embodiment, the sequence modeling module includes L sequence modeling units connected in sequence, and the output of the last sequence modeling unit is used as the input of the decoder;

[0015] The sequence modeling unit includes a multi-head attention mechanism and a visual Mamba module; the enhanced feature is input into the multi-head attention mechanism to obtain a first feature; the first feature is residually connected with the enhanced feature and subjected to layer normalization to obtain a second feature; the second feature is input into the visual Mamba module to obtain a third feature; the third feature is residually connected with the second feature and subjected to layer normalization to obtain a processed feature.

[0016] In one embodiment, the visual Mamba module includes a first Mamba branch, a second Mamba branch, and normalized attention;

[0017] The first Mamba branch includes a linear layer, a depth-separable convolutional layer, a SiLU activation function, a 2D-SSM layer, and a LayerNorm connected in sequence, and the second Mamba branch includes a linear layer and a SiLU activation function connected in sequence; the second feature is input to the first Mamba branch and the second Mamba branch respectively, and the outputs of the first Mamba branch and the second Mamba branch are Hadamard multiplied and added to the second feature to obtain an aggregated feature;

[0018] The attention module includes LayerNorm, convolutional layer and channel attention connected in sequence. The aggregated features are input into the normalized attention to obtain the attention features. The attention features are added to the aggregated features to obtain the third features.

[0019] In one embodiment, the decoder includes a first decoding branch and a second decoding branch; the first decoding branch includes a Mini-Unet, and the second decoding branch includes a position encoding module;

[0020] The processed features are input into the first decoding branch to obtain the semantic segmentation result;

[0021] Randomly generate a zero vector, and input the zero vector into the second decoding branch for position encoding to obtain a position encoding result;

[0022] The semantic segmentation result and the position encoding result are multiplied to obtain a first multiplication result;

[0023] The first multiplication result is multiplied by the processed features and input into the linear layer to obtain the output of the decoder.

[0024] In a second aspect, a scene text recognition device based on a hybrid architecture of visual Mamba and Transformer is provided, comprising:

[0025] The data set acquisition module is used to acquire a text image data set; the samples in the text image data set include scene text images in a variety of different fonts, different backgrounds and different noise conditions;

[0026] The network building module is used to build a scene text recognition network, which includes a spatial transformer network STN, ResNet, a feature enhancement module, a sequence modeling module, and a decoder;

[0027] The network training module is used to train the scene text recognition network based on the text image dataset to obtain the trained scene text recognition network; the training process includes:

[0028] The sample is input into the scene text recognition network, and the spatial transformer network STN performs geometric transformation correction on the sample to obtain the corrected image; ResNet extracts features from the corrected image to obtain local visual features; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features; the sequence modeling module uses visual Mamba and Transformer to capture the long-range dependencies and global context information in the scene text to obtain processed features for the enhanced features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probability of all text categories is selected as the recognition result of the character;

[0029] The recognition module is used to input the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character.

[0030] In a third aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the scene text recognition method based on the hybrid architecture of visual Mamba and Transformer is implemented.

[0031] In a fourth aspect, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned scene text recognition method based on the hybrid architecture of visual Mamba and Transformer.

[0032] Compared with the prior art, the present application has the following beneficial effects: the scene text recognition method based on the hybrid architecture of visual Mamba and Transformer of the present application utilizes visual Mamba to effectively compress and model the visual context, and is successfully used in fields such as visual prediction; combined with the multi-head attention mechanism of Transformer, the ability to perceive global context information is improved, the ability to model the sequence of text images is enhanced, and the accuracy of the text recognition algorithm is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The present application may be better understood by referring to the following description given in conjunction with the accompanying drawings, which together with the following detailed description are included in this specification and form a part of this specification. In the drawings:

[0034] Figure 1 A flowchart of a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer is shown;

[0035] Figure 2 shows a schematic diagram of a feature enhancement module;

[0036] Figure 3 shows a schematic diagram of a sequence modeling module;

[0037] Figure 4 shows a schematic diagram of the visual Mamba module;

[0038] Figure 5 shows a schematic diagram of a decoder;

[0039] Figure 6 A schematic diagram of the recognition results of some images is shown. DETAILED DESCRIPTION

[0040] The exemplary embodiments of the present application will be described below in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many implementation-specific decisions can be made in the process of developing any such actual embodiments in order to achieve the specific goals of the developer, and these decisions may vary from embodiment to embodiment.

[0041] It is also necessary to explain here that, in order to avoid obscuring the present application due to unnecessary details, only the device structure closely related to the scheme according to the present application is shown in the drawings, while other details that are not closely related to the present application are omitted.

[0042] It should be understood that the present application is not limited to the described implementation forms due to the following description with reference to the accompanying drawings. In this article, where feasible, the embodiments can be combined with each other, features between different embodiments can be replaced or borrowed, and one or more features can be omitted in one embodiment.

[0043] The present application embodiment provides a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer. Figure 1 A flowchart of a scene text recognition method based on a hybrid architecture of visual Mamba and Transformer is shown in FIG. Figure 1 , methods include:

[0044] Step S1, obtaining a text image dataset; the samples in the text image dataset include scene text images in a variety of different fonts, different backgrounds and different noise conditions.

[0045] Here, the text image datasets adopt the synthetic scene text image datasets MJSynth and SynthTex.

[0046] Before conducting network training based on the text image dataset, the samples can be preprocessed, the images can be scaled to a size of 32×128, and the images can be data enhanced by using methods such as random rotation, perspective transformation, and Gaussian noise addition to improve the robustness and generalization ability of the model.

[0047] Step S2, constructing a scene text recognition network, which includes a spatial transformer network STN (Spatial Transformer Network), a ResNet (Residual Network), a feature enhancement module, a sequence modeling module and a decoder.

[0048] Step S3, training the scene text recognition network based on the text image data set to obtain a trained scene text recognition network; the training process includes:

[0049] The sample is input into the scene text recognition network, and the spatial transformer network STN performs geometric transformation correction on the sample to obtain the corrected image; ResNet extracts features from the corrected image to obtain local visual features of 1 / 4 the size of the original image; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features, thereby improving the network's perception of the text area; the sequence modeling module uses visual Mamba and Transformer to capture the long-range dependencies and global context information in the scene text to obtain processed features for the enhanced features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character;

[0050] Here, the spatial transformer network STN will perform geometric transformation correction on the image to correct the image's tilt, rotation, and perspective distortion, reducing the difficulty of subsequent recognition.

[0051] After obtaining the character recognition results, the predicted labels are compared with the real labels, and the cross entropy loss function is used to calculate the loss. The loss is back-propagated to update the network parameters through the gradient descent method, thereby achieving end-to-end training. When the network loss value no longer fluctuates greatly and reaches a convergence state, the network training is terminated, and the network and weight files for scene text recognition are finally obtained.

[0052] Using the trained network and weights, the natural scene text image datasets ICDAR2013, IIIT5K, ICDAR2015, CUTE80, SVT, and SVTP are used as test sets, which contain text images from actual environments, to test and evaluate network performance and output recognition results.

[0053] The synthetic scene text dataset is used for training, combined with data enhancement techniques (such as random rotation, random cropping, geometric transformation, etc.) to further improve the generalization ability of the network. During the training process, the Adam optimizer and learning rate decay strategy are used to ensure network convergence and avoid overfitting. In the testing phase, the natural scene text dataset is used to verify the trained network and evaluate the text recognition accuracy and robustness of the model in real scenes.

[0054] Step S4, input the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, select the text corresponding to the maximum value of the corresponding probabilities of all text categories as the recognition result of the character.

[0055] In this embodiment, Visual Mamba is used to effectively compress and model the visual context, and is successfully used in fields such as visual prediction. Combined with the Transformer's multi-head attention mechanism, the ability to perceive global context information is improved, the ability to model the sequence of text images is enhanced, and the accuracy of the text recognition algorithm is effectively improved.

[0056] In one embodiment, Figure 2 A schematic diagram of a feature enhancement module is shown, the feature enhancement module comprising a first convolution branch, a second convolution branch, a third convolution branch and normalized attention;

[0057] The first convolution branch includes two convolution layers connected in sequence, and the second convolution branch has the same structure as the third convolution branch, including three convolution layers connected in sequence and a dilated convolution layer;

[0058] The first convolution branch, the second convolution branch, and the third convolution branch output features of different scales. The features of different scales are input into the convolution layer. The convolution results are then input into the normalized attention for weighted processing to obtain enhanced features.

[0059] In this embodiment, multiple convolution branches are used to extract features at different scales and directions to expand the local perception ability of the network, and normalized attention is used to weight the features of the multi-branch convolution output to highlight important features and suppress irrelevant or interfering information.

[0060] In one embodiment, Figure 3 A schematic diagram of a sequence modeling module is shown. The sequence modeling module includes L sequence modeling units connected in sequence, and the output of the last sequence modeling unit is used as the input of the decoder; here, the specific value of L is selected according to actual conditions, for example, it can be 5 or 7.

[0061] The sequence modeling unit includes a multi-head attention mechanism and a visual Mamba module; the enhanced feature is input into the multi-head attention mechanism to obtain a first feature; the first feature is residually connected with the enhanced feature and subjected to layer normalization to obtain a second feature; the second feature is input into the visual Mamba module to obtain a third feature; the third feature is residually connected with the second feature and subjected to layer normalization to obtain a processed feature.

[0062] In one embodiment, Figure 4 A schematic diagram of a visual Mamba module is shown, wherein the visual Mamba module includes a first Mamba branch, a second Mamba branch, and an attention module;

[0063] The first Mamba branch includes a linear layer, a depth-separable convolutional layer, a SiLU activation function, a 2D-SSM layer (2D selective scanning layer) and a layer normalization LayerNorm connected in sequence, and the second Mamba branch includes a linear layer and a SiLU activation function connected in sequence; the second feature is input to the first Mamba branch and the second Mamba branch respectively, and the outputs of the first Mamba branch and the second Mamba branch are Hadamard multiplied and added to the second feature to obtain an aggregated feature;

[0064] The attention module includes layer normalization LayerNorm, convolutional layer and channel attention connected in sequence. The aggregated features are input into the attention module to obtain the attention features; the attention features are added to the aggregated features to obtain the third feature.

[0065] In one embodiment, Figure 5 A schematic diagram of a decoder is shown, the decoder comprising a first decoding branch and a second decoding branch; the first decoding branch comprises a Mini-Unet, and the second decoding branch comprises a position encoding module;

[0066] The processed features are input into the first decoding branch to obtain the semantic segmentation result;

[0067] Randomly generate a zero vector, and input the zero vector into the second decoding branch for position encoding to obtain a position encoding result;

[0068] The semantic segmentation result and the position encoding result are multiplied to obtain a first multiplication result;

[0069] The first multiplication result is multiplied by the processed features and input into the linear layer to obtain the output of the decoder.

[0070] The scene text recognition method based on the visual Mamba and Transformer hybrid architecture of the aforementioned embodiment is used to obtain the recognition result of the text image. Figure 6 A schematic diagram of the recognition results of some images is shown.

[0071] Using the same inventive concept as the scene text recognition method based on the hybrid architecture of visual Mamba and Transformer, this embodiment also provides a scene text recognition device based on the hybrid architecture of visual Mamba and Transformer corresponding thereto, including:

[0072] The data set acquisition module is used to acquire a text image data set; the samples in the text image data set include scene text images in a variety of different fonts, different backgrounds and different noise conditions;

[0073] The network building module is used to build a scene text recognition network, which includes a spatial transformer network STN, ResNet, a feature enhancement module, a sequence modeling module, and a decoder;

[0074] The network training module is used to train the scene text recognition network based on the text image dataset to obtain the trained scene text recognition network; the training process includes:

[0075] The sample is input into the scene text recognition network, and the spatial transformer network STN performs geometric transformation correction on the sample to obtain the corrected image; ResNet extracts features from the corrected image to obtain local visual features; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features; the sequence modeling module uses visual Mamba and Transformer to capture the long-range dependencies and global context information in the scene text to obtain processed features for the enhanced features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probability of all text categories is selected as the recognition result of the character;

[0076] The recognition module is used to input the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character.

[0077] The scene text recognition device based on the hybrid architecture of visual Mamba and Transformer of this embodiment has the same inventive concept as the scene text recognition method based on the hybrid architecture of visual Mamba and Transformer mentioned above. Therefore, the specific implementation method of the device can be seen in the implementation example part of the scene text recognition method based on the hybrid architecture of visual Mamba and Transformer mentioned above, and its technical effect corresponds to the technical effect of the above method, which will not be repeated here.

[0078] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned scene text recognition method based on the hybrid architecture of visual Mamba and Transformer is implemented.

[0079] An embodiment of the present application provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the above-mentioned scene text recognition method based on the hybrid architecture of visual Mamba and Transformer is implemented.

[0080] The above are only various implementations of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A scene text recognition method based on a hybrid architecture of visual Mamba and Transformer, characterized in that: include: Get text image dataset; The samples in the text image dataset include scene text images in a variety of different fonts, different backgrounds and different noise conditions; Constructing a scene text recognition network, wherein the scene text recognition network includes a spatial transformer network STN, a ResNet, a feature enhancement module, a sequence modeling module and a decoder; Training the scene text recognition network based on the text image data set to obtain a trained scene text recognition network; The training process includes: The sample is input into the scene text recognition network, and the spatial transformation network STN performs geometric transformation correction on the sample to obtain a corrected image; the ResNet extracts features from the corrected image to obtain local visual features; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features; the sequence modeling module uses visual Mamba and Transformer to capture long-range dependencies and global context information in the scene text for the enhanced features to obtain processed features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character; Inputting the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, selecting the text corresponding to the maximum value of the corresponding probabilities of all text categories as the recognition result of the character; The sequence modeling module includes L sequence modeling units connected in sequence, and the output of the last sequence modeling unit is used as the input of the decoder; The sequence modeling unit includes a multi-head attention mechanism and a visual Mamba module; the enhanced feature is input into the multi-head attention mechanism to obtain a first feature; the first feature is residually connected with the enhanced feature and subjected to layer normalization processing to obtain a second feature; the second feature is input into the visual Mamba module to obtain a third feature; the third feature is residually connected with the second feature and subjected to layer normalization processing to obtain the processed feature.

2. The method according to claim 1, characterized in that The feature enhancement module includes a first convolution branch, a second convolution branch, a third convolution branch and an attention mechanism; The first convolution branch includes two convolution layers connected in sequence, and the second convolution branch has the same structure as the third convolution branch, including three convolution layers and a hole convolution layer connected in sequence; The first convolution branch, the second convolution branch, and the third convolution branch output features of different scales, and the features of different scales are input into the convolution layer. The convolution results are then input into the attention mechanism for weighted processing to obtain the enhanced features.

3. The method according to claim 1, characterized in that The visual Mamba module includes a first Mamba branch, a second Mamba branch and an attention module; The first Mamba branch includes a linear layer, a depth-separable convolutional layer, a SiLU activation function, a 2D-SSM layer and a LayerNorm connected in sequence, and the second Mamba branch includes a linear layer and a SiLU activation function connected in sequence; the second feature is input into the first Mamba branch and the second Mamba branch respectively, and the outputs of the first Mamba branch and the second Mamba branch are Hadamard multiplied and added to the second feature to obtain an aggregated feature; The attention module includes a LayerNorm, a convolutional layer and a channel attention layer connected in sequence, and the aggregated features are input into the attention module to obtain attention features; The attention feature is added to the aggregation feature to obtain the third feature.

4. The method according to claim 1, characterized in that The decoder includes a first decoding branch and a second decoding branch; the first decoding branch includes a Mini-Unet, and the second decoding branch includes a position encoding module; The processed features are input into the first decoding branch to obtain a semantic segmentation result; Randomly generate a zero vector, and input the zero vector into the second decoding branch for position coding to obtain a position coding result; Multiplying the semantic segmentation result and the position encoding result to obtain a first multiplication result; The first multiplication result is multiplied by the processed feature and input into a linear layer to obtain an output of a decoder.

5. A scene text recognition device based on a hybrid architecture of visual Mamba and Transformer, characterized in that: include: A data set acquisition module is used to acquire text image data sets; The samples in the text image dataset include scene text images in a variety of different fonts, different backgrounds and different noise conditions; A network construction module, used to construct a scene text recognition network, wherein the scene text recognition network includes a spatial transformation network STN, a ResNet, a feature enhancement module, a sequence modeling module and a decoder; A network training module, used for training the scene text recognition network based on the text image data set to obtain a trained scene text recognition network; The training process includes: The sample is input into the scene text recognition network, and the spatial transformation network STN performs geometric transformation correction on the sample to obtain a corrected image; the ResNet extracts features from the corrected image to obtain local visual features; the feature enhancement module uses a multi-branch convolution structure to capture features of different scales for the local visual features, and performs weighted processing on the features based on the attention mechanism to obtain enhanced features; the sequence modeling module uses visual Mamba and Transformer to capture long-range dependencies and global context information in the scene text for the enhanced features to obtain processed features; the decoder decodes the processed features to obtain the probability that each character in the sample belongs to each text category; for each character, the text corresponding to the maximum value of the corresponding probabilities of all text categories is selected as the recognition result of the character; A recognition module, used for inputting the scene text image to be recognized into the trained scene text recognition network to obtain the probability that each character in the scene text image to be recognized belongs to each text category; for each character, selecting the text corresponding to the maximum value of the corresponding probabilities of all text categories as the recognition result of the character; The sequence modeling module includes L sequence modeling units connected in sequence, and the output of the last sequence modeling unit is used as the input of the decoder; The sequence modeling unit includes a multi-head attention mechanism and a visual Mamba module; the enhanced feature is input into the multi-head attention mechanism to obtain a first feature; the first feature is residually connected with the enhanced feature and subjected to layer normalization processing to obtain a second feature; the second feature is input into the visual Mamba module to obtain a third feature; the third feature is residually connected with the second feature and subjected to layer normalization processing to obtain the processed feature.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the scene text recognition method based on the visual Mamba and Transformer hybrid architecture described in any one of claims 1 to 4 is implemented.

7. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements the scene text recognition method based on the visual Mamba and Transformer hybrid architecture as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Natural scene text recognition method based on geometric prior and knowledge graph

    CN114821609A

  • Deep hash image retrieval method based on frequency domain decoupling and visual Mamba

    CN118820508A