Scene text recognition method based on optimized multi-modal vision and language processing

By using a hybrid convolution-Transformer hybrid neural network and a multi-scale attention module in scene text recognition, combining learning position coding and self-mask decoder to enhance language feature modeling, a bidirectional multimodal interaction module is designed to achieve deep fusion, solving the high computational complexity and noise problems in the existing multimodal methods, and achieving efficient and accurate scene text recognition.

CN120182958APending Publication Date: 2025-06-20SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510333961.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing multimodal methods rely on iterative correction of language models in scene text recognition, resulting in high computational complexity and low efficiency, and inaccurate recognition of visual models introduces noise, affecting overall performance.

Method used

Mixed convolution-Transformer hybrid neural network is used as a visual model, combining multi-scale attention modules to extract visual features; learning position coding and self-mask decoder are introduced into the language model to enhance language feature modeling capabilities; designing a bidirectional multimodal interaction module to achieve deep fusion of visual and language features through a bidirectional cross attention mechanism.

Benefits of technology

It significantly improves the efficiency and accuracy of scene text recognition, reduces the dependence on iterative correction, achieves a good balance of performance and efficiency, and improves the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182958A_ABST
    Figure CN120182958A_ABST
Patent Text Reader

Abstract

The invention discloses a scene text recognition method based on optimized multi-modal vision and language processing, which comprises the following steps of: firstly, normalizing image data; and then inputting the preprocessed data into the optimized visual model. The visual model extracts multi-scale space and semantic features through a convolution-Transform hybrid neural network, and enhances the feature expression ability by using a multi-scale attention mechanism; the language model corrects character probability vectors output by the visual model, and learnable position codes are introduced to optimize representation of the features. Through designing a bidirectional multi-modal interaction module, visual and language features are fused, and a self-adaptive fusion mechanism is used to generate high-quality multi-modal joint feature representation. In the application stage, the optimized model is deployed through an efficient reasoning framework, and the speed and accuracy of scene text recognition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scene text recognition, and in particular to a scene text recognition method based on optimized multi-modal vision and language processing. Background Art

[0002] Scene text recognition is an important task in the field of computer vision, aiming to accurately recognize text information from images in the real environment, and is widely used in information processing, autonomous driving, digital document management and other fields. For example, it can help search engines, social media platforms and image databases index and classify images more efficiently; in autonomous driving, driving safety can be improved by recognizing traffic signs and road signs; in document management, printed documents or handwritten notes are converted into editable digital formats, thereby optimizing the storage and retrieval efficiency of documents. Generally speaking, scene text recognition has important value in building an intelligent and digital society and will continue to promote the innovation and application of technologies. Different from traditional optical character recognition technologies, scene text recognition needs to deal with more complex real-world scenarios, including text blur, occlusion, tilt, bending deformation, and illumination changes. These complexities not only increase the difficulty of research, but also pose higher requirements for the robustness and generalization ability of algorithms.

[0003] Previous methods mainly regarded scene text recognition as a visual task and only used visual models to solve it. This method completely relies on visual information and has poor performance when dealing with low-quality or occluded text images. Recently, multi-modal methods have achieved encouraging results, which can be divided into two categories: simple multi-modal fusion methods and multi-modal interaction methods. The simple multi-modal fusion method first uses a visual model to obtain the probability distribution on the character sequence, and then follows a language model to correct the probability distribution. Some existing multi-modal interaction methods enhance the information interaction between visual features and language features by inputting the concatenated features of the two modalities into a Transformer. However, the methods of these multi-modal models either have too simple a fusion method or lack sufficient interactivity, and without exception, they all adopt a strategy of iterative correction for the language model, and this iterative process usually brings significant time costs.

[0004] In multi-modal methods, the performance is mainly affected by three modules: the visual model, the language model, and the multi-modal fusion module. However, existing multi-modal methods usually rely too much on the language model and ignore the leading role of the visual model. The visual model is the core for extracting text image information and provides input for the language model at the same time. Therefore, its performance is crucial for the accuracy of scene text recognition. Since inaccurate recognition of the visual model often introduces noise, existing methods have to rely on the iterative correction strategy of the language model to obtain better results. However, this dependence not only increases the computational complexity but also limits the efficiency of the model. In contrast, improving the ability of the visual model can significantly reduce the noise input, thereby enhancing the overall performance. At the same time, a powerful language model can effectively correct multiple noise problems in a single correction process, thus reducing the need for iterative correction. When both the visual model and the language model can provide high-quality features, combining a powerful bidirectional multi-modal interaction module can further improve the recognition performance. The present invention can not only significantly improve the efficiency and accuracy of multi-modal methods but also effectively eliminate the dependence on iterative correction, providing a more efficient solution for constructing a real-time scene text recognition system. Summary of the Invention

[0005] The object of the present invention is to overcome the shortcomings and deficiencies of existing multi-modal methods, and propose a scene text recognition method based on optimizing multi-modal vision and language processing, which can greatly improve the recognition efficiency and also significantly improve the accuracy.

[0006] To achieve the above object, the technical solution provided by the present invention is: a scene text recognition method based on optimizing multi-modal vision and language processing, including the following steps:

[0007] S1: Obtain natural scene image data, perform normalization processing on the image data to obtain text pictures of the same size;

[0008] S2: Input the text picture into a pre-trained visual model, output the visual features of the text picture, and then obtain a character probability vector through linear prediction; wherein, the visual model is composed of a convolutional-Transformer hybrid neural network and a multi-scale attention module. The convolutional-Transformer hybrid neural network extracts visual features from the text picture, and the visual features include spatial information and semantic information. The multi-scale attention module is used to fuse the spatial information and semantic information of the visual features to alleviate the problem of attention drift;

[0009] S3: Input the character probability vector into a pre-trained language model to correct the character probability vector and output the corresponding language features. The language model consists of a learnable position encoding function and a self-masking decoder. The learnable position encoding function is used to enhance the adaptability of the language model to input data, and the self-masking decoder performs bidirectional prediction through a masked attention mechanism.

[0010] S4: Input the visual features and language features into a pre-trained bidirectional multimodal interaction module to obtain multimodal joint features. The bidirectional multimodal interaction module includes a self-attention mechanism, a residual addition and normalization module, a bidirectional cross-attention mechanism, a feed-forward neural network, and an adaptive fusion module. The self-attention mechanism is used to further extract visual and language features. The residual addition and normalization module is used to retain the original features and stabilize the training. The bidirectional cross-attention mechanism is used to achieve deep fusion of cross-modal features. The feed-forward neural network is used to further process the features. The adaptive fusion module is used to generate high-quality multimodal joint features.

[0011] S5: Map the multimodal joint features to a common representation space through a linear layer. Subsequently, for the linear layer output of the multimodal joint features, calculate the probability distribution of its character categories through the softmax function, and select the category with the highest probability as the text prediction result.

[0012] Furthermore, in step S2, the convolutional-Transformer hybrid neural network includes two convolutional modules and one Vision-Transformer module. For an image x with dimensions H×W×3, first capture the low-level spatial information of the visual features through two convolutional modules, and then pass through a Vision-Transformer module to capture the high-level semantic information of the visual features. Between each module, perform downsampling using a convolutional layer with a stride of 2 to reduce the feature map to half of its previous spatial resolution. After the image x is processed by two convolutional modules and one Vision-Transformer module, three feature maps with different scales will be obtained. and

[0013]

[0014] where and represent two convolutional modules, V B represents the Vision-Transformer module, P represents the process of first applying the downsampling operation and then applying the patch embedding function. Denote the real number space, \(H\) denote the image height, \(W\) denote the image width, \(C_1\), \(C_2\) and \(C_3\) denote the dimensions of each feature map;

[0015] Design a multi-scale attention module to further extract visual features and fuse the spatial and semantic information of visual features; the multi-scale attention module consists of one self-attention mechanism and two cross-attention mechanisms, and each layer of the attention mechanism is composed of different Query, Key and Value; in the first layer of the self-attention mechanism, the Query is set based on the fusion of the global features at the current scale and the position embedding; in the subsequent two layers of cross-attention mechanisms of the module, the output of the previous layer of the attention mechanism is also used to construct the Query of the current layer of the attention mechanism;

[0016] To obtain the global features at each scale, perform global average pooling operation on the feature maps at each scale, and then perform linear mapping; the acquisition process of the global features of the \(i\)-th layer is expressed as follows:

[0017]

[0018] where \(L\) R denotes the linear mapping function, and then perform reshaping operation, \(GAP\) denotes the global average pooling function, denotes the \(i\)-th feature map, \(O\) denotes the length of the character sequence, and \(C\) denotes the dimension size;

[0019] Use learnable position encoding to obtain the position embedding of the character order:

[0020]

[0021] where denotes the learnable position encoding function of the \(i\)-th layer, denotes the position information embedding function of the \(i\)-th layer;

[0022] The entire calculation process of the multi-scale attention module is expressed as follows:

[0023]

[0024]

[0025] where \(Q\) i , \(K\) Ti and \(V\) i respectively denote the transpose of the Query vector, the Key vector and the Value vector of the \(i\)-th layer, denotes the trainable weight of the \(i\)-th layer, \(L\) denotes the linear function, \(\sigma\) denotes the activation function, denotes the visual features of the \(i\)-th layer, represents the visual features of the (i - 1)-th layer, softmax represents the softmax function, d i represents the scaling factor of the i-th layer.

[0026] Furthermore, in step S3, the self-masked decoder consists of a masked attention mechanism. In the masked attention mechanism, the Query is generated based on a learnable position encoding function. Meanwhile, by integrating the character probability vector and position information of the visual model, the Key and Value of this attention mechanism are constructed. In addition, the masked attention mechanism masks the currently predicted position and makes predictions based on bidirectional context. The construction process of the above-mentioned masked attention mechanism is as follows:

[0027] Q = E l (O)

[0028] K = V = σ(L(V pred ) + L(E l (O)))

[0029]

[0030] In the formula, Q, K T and V respectively represent the Query vector, the transpose of the Key vector, and the Value vector. V pred represents the character probability vector output by the visual model, E l represents the learnable position encoding function of the sequence; M is a mask matrix, whose diagonal elements are negative infinity and the rest are zero, F l represents the language features, and d represents the scaling factor.

[0031] Furthermore, in step S4, first, the correlation between visual and language modality features is enhanced through fixed position encoding, and then key features are further extracted from each modality using the self-attention mechanism; then, the bidirectional cross-attention mechanism is used to promote the interaction and correlation between different modality features. Subsequently, the feed-forward neural network is used to further process the features and generate multi-modal visual features and multi-modal language features; the bidirectional cross-attention mechanism uses the features of one modality as the Query and the features of the other modality as the Key and Value; finally, the gate mechanism is used to adaptively fuse the features of the two modalities; the bidirectional multi-modal interaction module adopts residual connection and layer normalization to retain the original features and ensure the stability of the module training. The bidirectional multi-modal interaction module is as follows:

[0032]

[0033] F out = G ⊙ F mv + (1 - G)Fml

[0034] In the formula, MM represents the process of multimodal interaction, including self-attention mechanism, residual and layer normalization, bidirectional cross-attention mechanism, and feed-forward neural network, F v represents the visual features of the third layer F mv and F ml represent multimodal visual features and multimodal language features, W g represents learnable weights, [,] represents concatenation, σ represents the sigmoid function, F out represents the final multimodal joint features, G represents the learned weights, and ⊙ represents element-wise multiplication.

[0035] Furthermore, the specific operation steps of step S5 are as follows:

[0036] S51: Use a linear layer to perform a linear transformation on the multimodal joint features and map them to a common representation space;

[0037] S52: Calculate the probability distribution of character categories for the output of the linear layer of the multimodal joint features through the softmax function, which is used to generate the prediction results for subsequent text recognition;

[0038] S53: Based on the predicted probability distribution of the character categories, select the category with the highest probability as the prediction result of the corresponding feature, that is, the text character corresponding to the input text image.

[0039] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0040] 1. For the visual model, the present invention adopts a hybrid convolutional-Transformer hybrid neural network to initially extract visual features. This network combines the advantages of convolutional networks in capturing local spatial features and the capabilities of Transformers in capturing global semantic contexts, thereby achieving a comprehensive extraction of spatial and semantic information. In addition, the present invention designs a novel multi-scale attention module, which significantly improves the text recognition ability of the visual model in complex scenarios by integrating spatial and semantic features. Specifically, the multi-scale attention module can simultaneously focus on feature maps of different scales, combine global visual features and position embedding information, effectively alleviating the attention drift problem that often occurs in previous methods, especially showing stronger robustness when dealing with blurred, tilted, or small text samples. At the same time, while improving the performance, this visual model maintains a low computational cost, achieving a balance between efficiency and effectiveness.

[0041] 2. For the language model, the present invention introduces a learnable position encoding method into the language model. Compared with the traditional fixed position encoding, this method can dynamically learn the most suitable position information representation according to the input data features. This flexible way of position information modeling can better adapt to different types of data, thus significantly improving the generalization ability of the language model in various complex scenarios. In addition, the present invention also combines the learnable position encoding with a self-masking decoder, enabling the language model to perform more accurate feature modeling based on bidirectional context, thereby effectively improving the character sequence correction performance, especially showing stronger adaptability when dealing with low-quality inputs or noisy visual features.

[0042] 3. For the multimodal fusion module, the present invention designs a bidirectional multimodal interaction module, the core of which is to fully realize the information interaction between visual and language features by using the bidirectional cross-attention mechanism. Different from the traditional simple feature concatenation or unidirectional interaction strategy, the bidirectional cross-attention mechanism can not only extract the complementary information of visual features to language features, but also enable the language features to reverse-optimize the visual feature representation, thus realizing the deep fusion between the two modalities. In addition, this module further extracts key features through the self-attention mechanism and adaptively fuses the feature representations of the two modalities through the gating mechanism, thereby generating high-quality fusion representations. Such high-quality fusion representations greatly improve the model's ability to recognize complex scene texts, and better performance can be achieved for both regular and irregular texts.

[0043] In summary, through the comprehensive optimization and design of the visual model, language model, and bidirectional multimodal interaction module, the present invention significantly enhances the independent performance of each module, and realizes the performance improvement of the overall model through the deep cooperation of multiple modalities. Different from the previous multimodal methods that rely on the language model iterative correction strategy, the present invention eliminates the dependence on iterative correction by improving the capabilities of the visual model and language model and combining an efficient bidirectional multimodal interaction module. Experiments show that while improving the recognition accuracy, the present invention greatly improves the inference efficiency, achieving a good balance between performance and efficiency. Finally, the present invention reaches the state-of-the-art performance on multiple benchmark datasets, providing an efficient and reliable solution for the scene text recognition task. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a framework diagram of the method of the present invention.

[0045] Figure 2 It is a framework diagram of the visual model.

[0046] Figure 3 It is a framework diagram of the language model.

[0047] Figure 4It is a framework diagram of a two-way multimodal interaction module. Specific implementation manners

[0048] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0049] As Figure 1 shown, this embodiment discloses a scene text recognition method based on optimized multimodal vision and language processing. Given a text image, the vision model uses a convolutional-Transformer hybrid neural network and a multi-scale attention module to generate visual features, and then obtains a character probability vector through linear prediction. Subsequently, the character probability vector is input into the language model, and the language model uses a self-masked decoder with learnable position encoding for character correction. Finally, the visual features and language features are input into a two-way multimodal interaction module. The features of the two modalities interact through bidirectional cross-attention, and then are combined through an adaptive fusion module to obtain the final output. It includes the following steps:

[0050] 1) Acquisition and preprocessing of graphic data:

[0051] First, the image data of natural scene text is acquired, and the following preprocessing steps are performed on the data:

[0052] The input image is normalized to unify the size to ensure that the model can process image data with different resolutions.

[0053] 2) Extraction of visual features:

[0054] As Figure 2 shown, the convolutional-Transformer hybrid neural network is used to process the image data to extract image features of spatial and semantic information. Between the modules, convolution with a stride of 2 is used for downsampling to reduce the spatial resolution of the feature map to 1 / 2 of the original. Then, a multi-scale attention module is used to further extract visual features. The specific steps are as follows:

[0055] 2.1) Input an image x with a dimension of H×W×3. First, two convolutional modules are used to capture the low-level spatial information of visual features, and two feature maps with different scales are obtained and

[0056]

[0057] where and represent two convolutional modules, P represents the process of first applying a downsampling operation and then applying a patch embedding function, Denote the real number space, \(H\) represents the height of the image, \(W\) represents the image width, and \(C1\) and \(C2\) represent the dimensions of each feature map.

[0058] 2.2) Input the feature map output by the convolutional layer into the Vision-Transformer module to capture the high-level semantic information of visual features, and obtain the feature map

[0059]

[0060] Among them, \(V\) B represents the Vision-Transformer module, and \(C3\) represents the dimension of this feature map .

[0061] 2.3) There are three layers in the multi-scale attention module, and each layer involves an attention mechanism. In the first-layer self-attention mechanism, the Query is set based on the fusion of the global features at the current scale and the position embedding. In the subsequent two-layer cross-attention mechanisms of this module, the output of the previous layer is also used to construct the Query. In order to obtain the global features of each scale, a global average pooling operation is performed on the feature map of each scale, followed by a linear mapping. The acquisition process of the global features of the \(i\)-th layer can be expressed as follows:

[0062]

[0063] Among them, \(L\) R represents the linear mapping function, followed by a reshaping operation, \(GAP\) represents the global average pooling function, represents the \(i\)-th feature map, \(O\) represents the length of the character sequence, and \(C\) represents the dimension.

[0064] Use learnable position encoding to obtain the position embedding of the character order:

[0065]

[0066] Among them, represents the learnable position encoding function of the \(i\)-th layer, represents the position embedding at the embedding function of the \(i\)-th layer.

[0067] The entire calculation process of the multi-scale attention module can be expressed as follows:

[0068]

[0069] Among them, \(Q\) i , \(K\) Ti and \(V\) irespectively represent the Query vector, the transpose of the Key vector, and the Value vector of the i-th layer. represents the trainable weight of the i-th layer, L represents the linear function, and σ represents the activation function. represents the visual feature of the i-th layer. represents the visual feature of the (i - 1)-th layer, and softmax represents the softmax function, d i represents the scaling factor of the i-th layer.

[0070] 3) Extraction of language features:

[0071] As Figure 3 shown, the character probability vector output by the visual model is input into the language model, and further correction is performed using a self-masking decoder, where this self-masking decoder is a variant of the Transformer decoder and is mainly composed of a masked attention mechanism.

[0072] The Query of the masked attention mechanism in the self-masking decoder is generated based on a learnable position encoding function. At the same time, by integrating the character probability and position information of the visual model, the Key and Value of the masked attention mechanism in the self-masking decoder are obtained. In addition, the masked attention mechanism masks the current predicted position and makes predictions based on bidirectional context. The construction process of the above-mentioned masked attention mechanism is as follows:

[0073] Q = E l (O)

[0074] K = V = σ(L(V pred ) + L(E l (O)))

[0075]

[0076] where Q, K T and V respectively represent the Query vector, the transpose of the Key vector, and the Value vector, V pred represents the character probability vector output by the visual model, E l represents the learnable position encoding function of the sequence. M is the mask matrix, whose diagonal elements are negative infinity and the rest are zero, F l represents the language feature, and d represents the scaling factor.

[0077] 4) Multimodal interaction and fusion:

[0078] As Figure 4As shown, first, the correlation between visual and linguistic modality features is enhanced through fixed-position embedding, and then the self-attention mechanism is used to further extract key features from each modality. Then, the bidirectional cross-attention mechanism is used to promote the interaction and correlation between different modality features. Subsequently, a feed-forward neural network is used to further process the features and generate multi-modal visual features and multi-modal linguistic features. More specifically, the bidirectional cross-attention mechanism uses the features of one modality as Query, and the features of the other modality as Key and Value. Finally, a gate mechanism is used to adaptively fuse the features of the two modalities. The bidirectional multi-modal interaction module adopts residual connection and layer normalization to retain the original features and ensure the stability of the module training. This process can be expressed as follows:

[0079]

[0080] F out = G ⊙ F mv + (1 - G)F ml

[0081] where MM represents the process of multi-modal interaction, including self-attention mechanism, residual and layer normalization, bidirectional cross-attention mechanism, and feed-forward neural network. F v represents the visual features of the third layer F mv and F ml represent multi-modal visual features and multi-modal linguistic features. W g represents learnable weights. [,] represents concatenation. σ represents the sigmoid function. F out represents the final multi-modal joint features. G represents the learned weights. ⊙ represents element-wise multiplication.

[0082] 5) Output of text prediction result:

[0083] The multi-modal joint features are linearly transformed using a linear layer and mapped to a common representation space. Subsequently, the linear layer output of the multi-modal joint features is used to calculate the probability distribution of character categories through the softmax function. Based on the probability distribution, the category with the highest probability is selected as the predicted result of the text, that is, the text character corresponding to the input text image.

[0084] All models are trained on two large synthetic datasets, namely MJSynth and SynthText. Six datasets from different sources are used to construct the test dataset, including IIIT5k, SVT, IC13, IC15, SVTP, and CUTE. The numbers shown in the first row of Table 1 represent the number of images in each dataset.

[0085] To verify the effectiveness of the method of the present invention, accuracy and inference time are used as evaluation criteria. On the above six datasets, comparative analysis is respectively carried out with three multimodal methods, namely SRN, ABINet and LevOCR. Among them, SRN and ABINet are simple multimodal methods with fast inference speed but low accuracy, while LevOCR is a multimodal interaction method with high accuracy but slow speed. The experimental results are shown in Table 1. The experimental results can fully prove the effectiveness of the method of the present invention. Compared with the existing methods, our model shows commendable performance in terms of accuracy and inference speed.

[0086]

[0087] Experimental conclusion: The experimental results show that the proposed method performs excellently on the six benchmark datasets. At the same time, the inference time is 28.2 ms, which is close to ABINet, but the performance is much higher than that of ABINet. Since the iterative correction strategy is not used, the inference speed is significantly better than that of the LevOCR method, and the performance is also better than that of the LevOCR method. The method of the present invention not only shows strong robustness when dealing with complex samples, but also achieves a good balance between accuracy and inference efficiency, providing an efficient and reliable solution for the scene text recognition task.

[0088] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A scene text recognition method based on optimized multimodal vision and language processing, characterized in that: The following steps are involved: S1: Acquire natural scene image data, normalize the image data, and obtain text images of uniform size; S2: Input the text image into a pre-trained visual model, output the visual features of the text image, and then obtain the character probability vector through linear prediction; wherein the visual model is composed of a convolution-Transformer hybrid neural network and a multi-scale attention module, the convolution-Transformer hybrid neural network extracts visual features from the text image, the visual features include spatial information and semantic information, and the multi-scale attention module is used to fuse the spatial information and semantic information of the visual features to alleviate the problem of attention drift; S3: inputting the character probability vector into a pre-trained language model, correcting the character probability vector, and outputting corresponding language features; wherein the language model is composed of a learnable position encoding function and a self-masking decoder, wherein the learnable position encoding function is used to enhance the adaptability of the language model to the input data, and the self-masking decoder performs bidirectional prediction through a masked attention mechanism; S4: Input the visual features and language features into a pre-trained bidirectional multimodal interaction module to obtain multimodal joint features; wherein the bidirectional multimodal interaction module includes a self-attention mechanism, a residual addition and normalization module, a bidirectional cross-attention mechanism, a feedforward neural network and an adaptive fusion module, wherein the self-attention mechanism is used to further extract visual and language features, the residual addition and normalization module is used to retain the original features and stabilize the training, the bidirectional cross-attention mechanism is used to achieve deep fusion of cross-modal features, the feedforward neural network is used to further process features, and the adaptive fusion module is used to generate high-quality multimodal joint features; S5: The multimodal joint features are mapped to the common representation space through linear layers respectively. Then, for the linear layer output of the multimodal joint features, the probability distribution of its character category is calculated through the softmax function, and the category with the largest probability is selected as the text prediction result.

2. The scene text recognition method based on optimized multimodal vision and language processing according to claim 1, characterized in that: In step S2, the convolution-Transformer hybrid neural network includes two convolution modules and one Vision-Transformer module; for an image x with a dimension of H×W×3, the low-level spatial information of the visual features is first captured by two convolution modules, and then a Vision-Transformer module is used to capture the high-level semantic information of the visual features; between each module, a convolution with a stride of 2 is used for downsampling to reduce the feature map to half of its previous spatial resolution; after the image x is processed by two convolution modules and one Vision-Transformer module, three feature maps of different scales will be obtained and In the formula, and Represents two convolution modules, V B represents the Vision-Transformer module, P represents the process of applying the downsampling operation first and then the patch embedding function, represents the real number space, H represents the image height, W represents the image width, C1, C2 and C3 represent the dimensions of each feature map; A multi-scale attention module is designed to further extract visual features and fuse the spatial information and semantic information of the visual features; the multi-scale attention module consists of a layer of self-attention mechanism and two layers of cross-attention mechanism, and each layer of attention mechanism consists of different queries, keys and values; in the first layer of self-attention mechanism, the query is set based on the fusion of the global features of the current scale and the position embedding; in the subsequent two layers of cross-attention mechanism of this module, the output of the previous layer of attention mechanism is also used to construct the query of the current layer of attention mechanism; In order to obtain the global features of each scale, a global average pooling operation is performed on the feature map of each scale, followed by linear mapping; the global features of the i-th layer The acquisition process is as follows: Where, L R represents a linear mapping function, followed by a reshaping operation, GAP represents a global average pooling function, represents the i-th feature map, O represents the length of the character sequence, and C represents the dimension size; Use learnable positional encoding to obtain positional embedding of character sequence: In the formula, represents the learnable positional encoding function of the i-th layer, represents the position information embedding function of the i-th layer; The entire calculation process of the multi-scale attention module is expressed as follows: In the formula, Q i , K Ti and V i Represent the Query vector, the transposed Key vector and the Value vector of the i-th layer respectively. represents the trainable weight of the i-th layer, L represents the linear function, σ represents the activation function, represents the visual features of the i-th layer, represents the visual features of the i-1th layer, softmax represents the softmax function, and d i represents the scaling factor of the i-th layer.

3. The scene text recognition method based on optimized multimodal vision and language processing according to claim 2, characterized in that: In step S3, the self-masking decoder is composed of a masked attention mechanism. In the masked attention mechanism, the query is generated based on a learnable position encoding function. At the same time, the Key and Value of the attention mechanism are constructed by integrating the character probability vector and position information of the visual model. In addition, the masked attention mechanism masks the current predicted position and predicts based on the bidirectional context. The construction process of the masked attention mechanism mentioned above is expressed as follows: Q=E l (O) K=V=σ(L(V pred )+L(E l (OR))) In the formula, Q, K T and V represent the Query vector, the transposed Key vector, and the Value vector, respectively. pred Represents the character probability vector output by the visual model, E l represents a learnable positional encoding function for a sequence; M is a mask matrix whose diagonal elements are negative infinity and the rest are zero, and F l represents the language feature, and d represents the scaling factor.

4. The scene text recognition method based on optimized multimodal vision and language processing according to claim 3 is characterized in that: In step S4, first, the correlation between the two modal features of vision and language is enhanced by fixed position encoding, and then the self-attention mechanism is used to further extract key features from each modality; then, the bidirectional cross attention mechanism is used to promote the interaction and correlation between the features of different modalities, and then a feedforward neural network is used to further process the features and generate multimodal visual features and multimodal language features; the bidirectional cross attention mechanism uses the features of one modality as the query and the features of the other modality as the key and value; finally, the gate mechanism is used to adaptively fuse the features of the two modalities; the bidirectional multimodal interaction module adopts residual connection and layer normalization to retain the original features and ensure the stability of the module training, and the bidirectional multimodal interaction module is expressed as follows: F out =G⊙F mv +(1-G)F ml Where MM represents the process of multimodal interaction, including self-attention mechanism, residual and layer normalization, bidirectional cross attention mechanism and feedforward neural network, F v Represents the visual features of the third layer F mv and F ml Represents multimodal visual features and multimodal language features, W g represents the learnable weight, [,] represents the connection, σ represents the sigmoid function, F out represents the final multimodal joint feature, G represents the learned weight, and ⊙ represents element-wise multiplication.

5. The scene text recognition method based on optimized multimodal vision and language processing according to claim 4, characterized in that: The specific operation steps of step S5 are as follows: S51: Use linear layers to perform linear transformation on multimodal joint features and map them to a common representation space; S52: Calculate the probability distribution of character categories through the softmax function for the linear layer output of the multimodal joint features, so as to generate the prediction results of subsequent text recognition; S53: Based on the predicted probability distribution of the character categories, the category with the largest probability is selected as the prediction result of the corresponding feature, that is, the text character corresponding to the input text image.

Citation Information

Cited By

  • Scene text recognition method and device based on multi-modal large language model

    CN120808329A

  • Double-branch diffusion three-dimensional scene generation method based on multi-modal semantic graph

    CN120833445A

  • Robot control method and device based on physical constraint embedding, equipment and medium

    CN120862691A

  • Abnormal behavior analysis and prediction system and method based on unmanned aerial vehicle monitoring system

    CN121121575A

  • Layered cross-domain information injection method and system for auxiliary driving system

    CN121659248A