Text recognition method, electronic device, and storage medium

The initial image features are acquired and position coded through the text recognition model, which solves the problem of noise influence in curved text recognition by multi-stage methods, and improves the accuracy of text recognition.

CN119888711BActive Publication Date: 2025-07-11HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510341770.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

When the existing multi-stage text recognition method deals with curved text, the cropping stage will introduce a lot of background noise, affecting the accuracy of character recognition.

Method used

The text recognition model is used to obtain the initial image features, and the position query vector is obtained through position encoding, combining feature information of different scales and positions to improve the accuracy of text recognition results.

Benefits of technology

The recognition accuracy of each area type, coordinate information and character information in the image to be identified is improved, especially the recognition effect of curved text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888711B_ABST
    Figure CN119888711B_ABST
Patent Text Reader

Abstract

The present application discloses a text recognition method, an electronic device, and a storage medium. The text recognition method includes: obtaining at least one initial image feature by using a text recognition model, where each initial image feature is obtained by performing feature extraction on an image to be recognized; performing position encoding on each initial image feature to obtain a position query vector corresponding to each initial image feature; and determining a text recognition result of the image to be recognized based on each initial image feature and each position query vector, where the text recognition result includes at least one of the following: the type of each region in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized. The above solution can improve the accuracy of the text recognition result obtained by text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a text recognition method, an electronic device, and a storage medium. Background Art

[0002] Scene text recognition mainly includes two tasks: text detection and recognition in images. Currently, the commonly used recognition technologies are implemented based on deep learning methods, and according to the processing process, they can be multi-stage methods. The multi-stage method usually first uses a detection algorithm to output the bounding rectangles of all texts in the image, crops each text box area, corrects and enhances it, and then uses a character recognition algorithm to recognize each character in the text. Among them, the multi-stage method improves the accuracy of regular text recognition by correcting a single text image, but for curved text, it will introduce a large amount of background noise in the cropping stage, affecting the accuracy of subsequent character recognition.

[0003] In view of the existing technical deficiencies, how to provide an effective text recognition solution is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] This application provides at least a text recognition method, an electronic device, and a storage medium.

[0005] This application provides a text recognition method, including: obtaining at least one initial image feature by using a text recognition model, where each initial image feature is obtained by performing feature extraction on an image to be recognized; performing position encoding on each initial image feature to obtain a position query vector corresponding to each initial image feature; and determining a text recognition result of the image to be recognized based on each initial image feature and each position query vector, where the text recognition result includes at least one of the following: the type of each region in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized.

[0006] This application provides a text recognition device, including: an obtaining module, an encoding module, and a determining module; the obtaining module is configured to obtain at least one initial image feature by using a text recognition model, where each initial image feature is obtained by performing feature extraction on an image to be recognized; the encoding module is configured to perform position encoding on each initial image feature to obtain a position query vector corresponding to each initial image feature; and the determining module is configured to determine a text recognition result of the image to be recognized based on each initial image feature and each position query vector, where the text recognition result includes at least one of the following: the type of each region in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized.

[0007] The present application provides an electronic device, including a memory and a processor. The processor is configured to execute program instructions stored in the memory to implement the above-mentioned text recognition method.

[0008] The present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the above-mentioned text recognition method is implemented.

[0009] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed based on each initial image feature to obtain a position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings

[0011] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.

[0012] Figure 1 is a flowchart of an embodiment of the text recognition method of the present application Figure One ;

[0013] Figure 2 is a flowchart of an embodiment of the text recognition method of the present application Figure Two ;

[0014] Figure 3 is a flowchart of an embodiment of the text recognition method of the present application Figure Three ;

[0015] Figure 4 is a flowchart of an embodiment of the text recognition method of the present application Figure Four ;

[0016] Figure 5 is a flowchart of an embodiment of the text recognition method of the present application Figure Five ;

[0017] Figure 6 is a flowchart of an embodiment of the text recognition method of the present application Figure Six ;

[0018] Figure 7 is a flowchart of an embodiment of the text recognition method of the present application Figure Seven ;

[0019] Figure 8aIt is a schematic diagram of a preset decoding layer in an embodiment of the text recognition method of the present application;

[0020] Figure 8b It is a schematic diagram of a sample image in an embodiment of the text recognition method of the present application;

[0021] Figure 8c It is a schematic diagram of a training module in an embodiment of the text recognition method of the present application;

[0022] Figure 8d It is a schematic diagram of a visual query vector construction module in an embodiment of the text recognition method of the present application;

[0023] Figure 8e It is a schematic diagram of a text-guided query vector construction module in an embodiment of the text recognition method of the present application;

[0024] Figure 8f It is a schematic diagram of a character-guided query vector construction module in an embodiment of the text recognition method of the present application;

[0025] Figure 9 It is a schematic diagram of the structure of an embodiment of the text recognition device of the present application;

[0026] Figure 10 It is a schematic diagram of the structure of an embodiment of the electronic device of the present application;

[0027] Figure 11 It is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0028] The following will describe the solutions of the embodiments of the present application in detail with reference to the accompanying drawings of the specification.

[0029] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0030] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0031] This application provides some text recognition methods and text recognition devices. The application scenarios of the text recognition method include text recognition of images with text in natural scenes. The execution subject of the text recognition method can be any device with text recognition capabilities, such as a text recognition device. For example, the text recognition device can be set in a terminal device, a server, or other processing devices. Among them, the terminal device can be a device for text recognition, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementation manners, the text recognition method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0032] Please refer to Figure 1 , Figure 1 which is a flowchart of an embodiment of the text recognition method of this application. Figure One Specifically, the text recognition method may include the following steps:

[0033] Step S11: Obtain at least one initial image feature by using a text recognition model.

[0034] Each initial image feature is obtained by performing feature extraction on the image to be recognized. Each initial image feature can be obtained by performing feature extraction on the image to be recognized in a natural scene. Among them, the image to be recognized can be an image containing the text to be recognized. Specifically, the image to be recognized can be a natural scene picture, a document scan, a handwritten font picture, a screenshot, an emoji with text information, a design work, a medical image data, etc. Among them, the natural scene picture can be a picture with text information such as a street sign, a billboard, a store sign, a road sign, a menu, a product package, etc. Among them, the document scan can be a digital scan of a document such as a contract, a file, a report, an invoice, a business card, etc., and the text information in it needs to be extracted. Among them, the handwritten font picture can be a handwritten text picture with different styles and qualities, including handwritten notes, letters, envelopes, business cards, calligraphy works, etc. Among them, the screenshot can be a screenshot picture containing text information such as an application program interface, a web page screenshot, a mobile phone page screenshot, etc. Among them, the art or design work can be the text information in art or design works such as posters, advertising design drafts, book covers, paintings, etc. Among them, the medical image data can be pictures in the medical field such as medical reports, imaging diagnosis reports, annotation information in medical images, etc. This is only for illustration and does not limit the type of the image to be recognized.

[0035] The text recognition model can be a model with text recognition capabilities deployed on a text recognition device. Among them, the input of the text recognition model can be the image to be recognized, and the output of the text recognition model can be the text recognition result of the image to be recognized. It can be understood that each image to be recognized can correspond to at least one initial image feature. Each image to be recognized corresponds to one text recognition result.

[0036] In some application scenarios, the method of feature extraction for the image to be recognized can be to input the image to be recognized into the backbone network. The choice of the backbone network is relatively flexible and can be any backbone network with strong feature extraction capabilities. For example, ResNet series, Swin Transformer series, etc. In some other application scenarios, the method of feature extraction for the image to be recognized can be to input the image to be recognized into the backbone network and the feature encoder in sequence, so as to learn the global image information of the image to be recognized, adjust the attention of features at different positions in the image to be recognized, strengthen the key features, and weaken the irrelevant features, so as to obtain at least one initial image feature. In some application scenarios, the initial image feature can include the image feature obtained by feature extraction of the image to be recognized. The at least one initial image feature can be one initial image feature corresponding to the image to be recognized, or multiple initial image features corresponding to the image to be recognized. It can be understood that at least one means one or more.

[0037] Step S12: Perform position encoding on each initial image feature to obtain a position query vector corresponding to each initial image feature.

[0038] Each initial image feature can correspond to a feature vector. Each position query vector can be a feature vector that has richer feature information and image position information compared to the corresponding initial image feature.

[0039] The positional encoding can be to encode each initial image feature according to a preset encoding method. In some application scenarios, the preset encoding method can be to directly perform a preset positional encoding on each initial image feature and use the initial image feature after the preset positional encoding as the positional query vector corresponding to the initial image feature. Among them, the preset positional encoding can be cosine positional encoding or sine positional encoding. In other application scenarios, the preset encoding method can also be to perform a preset feature extraction on the initial image feature after the preset positional encoding to obtain the positional query vector corresponding to the initial image feature. Among them, the preset feature extraction can be to input the initial image feature after the preset positional encoding into a preset feature extraction module. The preset feature extraction module is used to perform feature extraction on the input features. Among them, a preset feature extraction network can be set on the preset feature extraction module. Specifically, the preset feature extraction network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, an attention mechanism, or other feature extraction networks. Exemplarily, the preset feature extraction network in each prediction module can be a long short-term memory network (LSTM) in a recurrent neural network. In other application scenarios, the preset feature extraction network can be a preset perceptron network. Among them, the preset perceptron network can be obtained by at least one perceptron module. Each perceptron module in the at least one perceptron module can be a multi-layer perceptron (MLP). Each initial image feature can be aggregated into an initial image feature set. The above step S12 can be to use each initial image feature or the initial image feature set as the input of the preset perceptron network to obtain the positional query vectors of each initial image feature output by the preset perceptron network.

[0040] Step S13: Based on each initial image feature and each positional query vector, determine the text recognition result of the image to be recognized.

[0041] The text recognition result includes at least one of the following: the type of each region in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized.

[0042] Each region in the image to be recognized may be an image region corresponding to at least one image patch in the image to be recognized. The type of each region in the image to be recognized includes a foreground type or a background type. The foreground region may represent the image region where the text object is located. The background region may represent the image region where the non-text object is located. The coordinate information of each region in the image to be recognized may characterize the position information of the region relative to the image to be recognized. The character information of each region in the image to be recognized may characterize the prediction result of the text object corresponding to the region. It can be understood that when the type of any region in the image to be recognized is the foreground type, the character information of the region in the image to be recognized may characterize the prediction result of the text object corresponding to the region. When the type of any region in the image to be recognized is the background type, the character information of the region in the image to be recognized may characterize the preset text of the non-text object corresponding to the region. Wherein, the preset text may be a value of 0.

[0043] Optionally, a text recognition model may be pre-trained, using each initial image feature and each position query vector as the input of the text recognition model, and the output of the text recognition model is the text recognition result of the image to be recognized. Or post-process the output of the text recognition model to obtain the text recognition result of the image to be recognized. The post-processing may be to select the text recognition result with the highest confidence from the text recognition results corresponding to each initial image feature as the text recognition result of the image to be recognized.

[0044] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed based on each initial image feature to obtain the position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0045] In some embodiments, the above step S11 may include the following steps: First, use a text recognition model to obtain the image to be recognized. Then, perform multi-scale feature extraction on the image to be recognized to obtain a multi-scale feature set corresponding to the image to be recognized. The multi-scale feature set contains at least two scale features and the scales of each scale feature are different. Then, fuse the scale features in the multi-scale feature set to obtain at least one initial image feature.

[0046] The way to obtain the image to be recognized may be to collect the image to be recognized in real time through an image acquisition device, or to use the image to be recognized stored in a database, or to intercept a frame of the video file.

[0047] After obtaining the image to be recognized, the image to be recognized is used as the input of the backbone network, and multi-scale feature extraction of the image to be recognized is implemented in the backbone network to obtain a multi-scale feature set corresponding to the image to be recognized output by the backbone network. The multi-scale feature set contains at least two scale features and the scales of the scale features are different.

[0048] In some application scenarios, the above method of performing feature fusion on each scale feature in the multi-scale feature set to obtain at least one initial image feature can be to directly combine each scale feature in the scale feature set arbitrarily to obtain at least one initial image feature with the same number as the number of scale features. In some other application scenarios, the above method of performing feature fusion on each scale feature in the multi-scale feature set to obtain at least one initial image feature can be to splice each scale feature in the scale feature set to obtain at least one initial image feature with the same number as the number of scale features. In some other application scenarios, the above method of performing feature fusion on each scale feature in the multi-scale feature set to obtain at least one initial image feature can be to set weights for each scale feature according to the distance from the current scale feature, and then for each current scale feature, based on the weights of each scale feature, fuse all scale features to obtain the initial image feature corresponding to the current scale feature. In some other application scenarios, the above method of performing feature fusion on each scale feature in the multi-scale feature set to obtain at least one initial image feature can be to use each scale feature in the scale feature set as the input of a preset feature encoder to obtain the initial image features corresponding to each scale feature output by the preset feature encoder.

[0049] Exemplarily, after using the backbone network to perform multi-scale feature extraction on the image to be recognized to obtain multiple scale features, the above method of performing feature fusion on each scale feature in the multi-scale feature set to obtain at least one initial image feature can be: unify the number of channels of the multiple scale features by performing a convolution operation on each scale feature in the scale feature set to obtain each processed scale feature, and flatten and splice each processed scale feature to obtain the spliced feature. The spliced feature can be expressed as . In the spliced feature, B can represent the batch size, S represents the sum of the product of the width and height of each scale feature, and C represents the number of channels of the spliced feature. Use the spliced feature as the input of a preset feature encoder to obtain the initial image features corresponding to each scale feature output by the preset feature encoder. Each initial image feature output by the preset feature encoder can be expressed as . Each initial image feature In, B can represent the batch size, S represents the sum of the product of the width and height of each scale feature, and C represents the number of channels of each initial image feature. The preset feature encoder can be a Transformer feature encoder.

[0050] It can be understood that by sending the spliced feature F into the Transformer feature encoder, the global information of the image to be recognized can be learned, the attention of features at different positions in the spliced feature F can be adjusted, key features can be strengthened, and irrelevant features can be weakened, so as to obtain each initial image feature. Furthermore, each initial image feature has global context information, making the text recognition result obtained by using each initial image feature more comprehensive and accurate in subsequent processing.

[0051] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the text recognition method of the present application Figure Two .

[0052] In some embodiments, the above step S12 may include the following steps: For each initial image feature, execute the following steps as Figure 2 shown: Step S21: Input the initial image feature into the first perceptron module to obtain the sampling point coordinate information of the initial image feature output by the first perceptron module. Step S22: Use the preset-encoded sampling point coordinate information as the input of the second perceptron module to obtain the position query vector corresponding to the initial image feature output by the second perceptron module.

[0053] The first perception module and the second perceptron module may be multi-layer perceptrons (MLPs). Among them, the first perception module and the second perceptron module may have the same structure but different parameters.

[0054] The sampling point coordinate information of each initial image feature may be the position information of the preset number of sampling points of the Bezier curve related to the image to be recognized in the image to be recognized. The preset encoding process may be the above-mentioned preset position encoding. Specifically, the preset position encoding may be cosine position encoding or sine position encoding.

[0055] In some application scenarios, after step S21, the sampling point coordinate information of the initial image feature is subjected to a preset encoding process to obtain the preset-encoded sampling point coordinate information. In other application scenarios, before the above step S21, the initial image feature is input into a preset linear layer to obtain the confidence corresponding to the initial image feature. Based on the confidence ranking of the initial image features, at least one initial image feature that meets the preset ranking is selected from each initial image feature as the target initial image feature. The preset ranking may be the top k in the confidence ranking of the initial image features. Among them, the confidence ranking of the initial image features is obtained from the confidence corresponding to each initial image feature. Then, the above step S21 may be for each target initial image feature, input the target initial image feature into the first perceptron module to obtain the sampling point coordinate information of the initial image feature corresponding to the target initial image feature output by the first perceptron module.

[0056] In some other application scenarios, the above step S22 may be to use the sampled point coordinate information processed by preset encoding as the input of the second perceptron module to obtain the initial position query vector corresponding to the initial image features output by the second perceptron module. Obtain the content query vector. The content query vector can be randomly initialized or be a preset query vector. It can be understood that the scale of the content query vector and the scale of the initial position query vector can be the same scale, or the scale of the initial position query vector is higher than the scale of the content query vector. Fuse the initial position query vector corresponding to the above initial image features with the content query vector to obtain the position query vector corresponding to the initial image features.

[0057] Exemplarily, constructing the position query vector corresponding to each initial image feature may be composed of the initial position query vector and the content query vector added together. Among them, each initial image feature outputs the probability score of each initial image feature through the classification linear layer Linear, and obtains the index of the top k features ranked from high to low according to the score. Each initial image feature inputs the L sampled point coordinates of the central Bezier curve corresponding to each text instance into the Bezier coordinate converter composed of the first perceptron module MLP, where L refers to the unified text length. Obtain the L sampled point coordinates of the central Bezier curve corresponding to the features ranked top k according to the top k indexes , and use the sampled point coordinates as the sampled point coordinate information of the initial image features. After the sampled point coordinate information is encoded by cosine position encoding and input into the position encoder composed of the second perceptron module, the initial position query vector is obtained. Adding the position query vector and the content query vector gives the position query vector corresponding to each initial image feature in the above step S22 . The content query vector is randomly initialized.

[0058] In some application scenarios, the above step S13 may be to perform feature fusion on each initial image feature and each position query vector to obtain the target fusion feature corresponding to each initial image feature. Then, based on each target fusion feature, determine the text recognition result of the image to be recognized.

[0059] Please refer to Figure 3 , Figure 3 which is the flowchart of an embodiment of the text recognition method of the present application Figure Three .

[0060] In some embodiments, the above step S13 may include the following steps: Step S31: Input each initial image feature and each position query vector into a multi-layer decoder for feature fusion to obtain the target fusion feature corresponding to each initial image feature. Step S32: Based on each target fusion feature, determine the text recognition result of the image to be recognized.

[0061] The multi-layer decoder may be a decoder having at least two preset feature extraction networks. For example, the multi-layer decoder may have at least two attention mechanisms. The target fusion feature corresponding to each initial image feature includes the feature information combined between the initial image feature and the position query vector corresponding to the initial image feature.

[0062] The above step S13 may be to input each initial image feature and each position query vector into a multi-layer decoder to obtain the target fusion feature corresponding to each initial image feature output by the multi-layer decoder.

[0063] Please refer to Figure 4 , Figure 4 which is a schematic flow of an embodiment of the text recognition method of the present application Figure Four .

[0064] In some embodiments, the multi-layer decoder includes a plurality of preset decoding layers arranged in cascade. The structures of the preset decoding layers are the same but the parameters are different. The feature output by the last preset decoding layer is used as the feature output by the multi-layer decoder. The feature output by the multi-layer decoder is the target fusion feature corresponding to each initial image feature. Each preset decoding layer includes an intra-group attention module. The above step S31 may include the following steps: Step S41: Use the first query-key-value pair as the input of the intra-group attention module in the current preset decoding layer to obtain the first feature output by the intra-group attention module in the current preset decoding layer. The query data, key data, and value data in the first query-key-value pair are respectively obtained from each position query vector or from the feature output by the previous preset decoding layer. The current preset decoding layer is one of the plurality of preset decoding layers. Step S42: Fuse the first feature output by the intra-group attention module in the current preset decoding layer with the input of the intra-group attention module in the current preset decoding layer to obtain the first fusion feature. Step S43: Perform advanced fusion processing on the first fusion feature and each initial image feature to obtain the feature output by the current preset decoding layer.

[0065] The multi-layer decoder may include a number of preset decoding layers arranged in cascade. Each preset decoding layer has the same structure but different parameters, and the features output by the last preset decoding layer are used as the features output by the multi-layer decoder. In some application scenarios, each preset decoding layer may include a feature fusion module that can directly fuse the initial image features and the position query vectors. In other application scenarios, each preset decoding layer includes a cross-attention module. Among them, the cross-attention module included in each preset decoding layer can fuse the initial image features and the position query vectors to obtain the target fusion features corresponding to the initial image features.

[0066] In other application scenarios, each preset decoding layer may include a single cross-attention module, or may include other attention modules in addition to the cross-attention module. When the preset decoding layer includes a cross-attention module and other attention modules, the cross-attention module and other attention modules can be arranged serially or in parallel. The setting order of the cross-attention module and other attention modules in each preset decoding layer is not limited here.

[0067] It can be understood that the first features output by the intra-group attention module in the current preset decoding layer include the first features corresponding to the initial image features. The first fusion features include the first fusion features corresponding to the initial image features obtained by fusing the first features corresponding to the initial image features and the features input to the intra-group attention module in the current preset decoding layer.

[0068] In some application scenarios, the position query vectors or the features output by the previous preset decoding layer are directly used as the query data, key data, and value data in the first query key-value pair. In other application scenarios, the position query vectors or the features output by the previous preset decoding layer are subjected to a mapping process to obtain the mapped position query vectors or the mapped features output by the previous preset decoding layer. And the mapped position query vectors or the mapped features output by the previous preset decoding layer are used as the query data, key data, and value data in the first query key-value pair of the current preset decoding layer. The mapping process can be specifically implemented by different linear transformations. Thus, it is ensured that the query data, key data, and value data in the first query key-value pair are in the same vector space.

[0069] In some application scenarios, when the current preset decoding layer is the first preset decoding layer in a multi-layer decoder, the input to the intra-group attention module in the current preset decoding layer is the first query-key-value pair. Among them, the query data, key data, and value data in the first query-key-value pair are obtained from the position query vectors respectively. The first feature output by the intra-group attention module in the current preset decoding layer is fused with the input to the intra-group attention module in the current preset decoding layer, that is, the first feature output by the intra-group attention module in the current preset decoding layer is fused with the position query vectors to obtain the first fused feature. The specific fusion processing method can be cross-layer addition processing. The first feature output by the intra-group attention module in the current preset decoding layer and the position query vectors are used as the input to the first preset interaction module to obtain the first fused feature output by the first preset interaction module. Among them, the first preset interaction module can be a network module or algorithm module provided with a residual connection or skip connection.

[0070] In other application scenarios, when the current preset decoding layer is a non-first preset decoding layer in a multi-layer decoder, the input to the intra-group attention module in the current preset decoding layer is the first query-key-value pair. Among them, the query data, key data, and value data in the first query-key-value pair are obtained from the features output by the previous preset decoding layer respectively. The first feature output by the intra-group attention module in the current preset decoding layer is fused with the input to the intra-group attention module in the current preset decoding layer, that is, the first feature output by the intra-group attention module in the current preset decoding layer is fused with the features output by the previous preset decoding layer to obtain the first fused feature. The specific fusion processing method can be cross-layer addition processing. The first feature output by the intra-group attention module in the current preset decoding layer and the features output by the previous preset decoding layer are used as the input to the first preset interaction module to obtain the first fused feature output by the first preset interaction module.

[0071] The advanced fusion processing can be cross-fusion processing of at least two features. Specifically, the advanced fusion processing can be inputting the first fused feature and the initial image features into the cross-attention module in each preset decoding layer to obtain the features output by the cross-attention module.

[0072] In some application scenarios, step S43 above can be inputting the first fused feature and the initial image features into the cross-attention module in the current preset decoding layer for cross-attention processing to obtain the features output by the cross-attention module, and using the features output by the cross-attention module as the features output by the current preset decoding layer.

[0073] It can be understood that by first processing the position query vector through the intra-group attention module, the subsequent obtained first fusion feature can be further fused with each initial image feature, so that the feature output by the multi-layer decoder is a feature with richer feature information, thereby improving the accuracy of the subsequent text recognition result.

[0074] Please refer to Figure 5 , Figure 5 which is a flowchart of an embodiment of the text recognition method of the present application. Figure Five .

[0075] In some embodiments, each preset decoding layer includes a cross-attention module, and the above step S43 may include the following steps: Step S51: Use the second query key-value pair as the input of the cross-attention module in the current preset decoding layer to obtain the second feature output by the cross-attention module in the current preset decoding layer. The query data in the second query key-value pair is obtained from the first fusion feature, and the key data and value data in the second query key-value pair are obtained from each initial image feature respectively. Step S52: Fuse the second feature output by the cross-attention module in the current preset decoding layer with the first fusion feature to obtain a second fusion feature. Step S53: Based on the second fusion feature, obtain the feature output by the current preset decoding layer.

[0076] In some application scenarios, directly use the first fusion feature as the query data in the second query key-value pair, and directly use all the initial image features as the key data and value data in the second query key-value pair respectively. In other application scenarios, perform a mapping process on the first fusion feature to obtain the mapped first fusion feature. And, perform a mapping process on all the initial image features to obtain the mapped initial image features respectively. And use the mapped initial image features as the key data and value data in the second query key-value pair, and use the mapped first fusion feature as the query data in the second query key-value pair. The mapping process can be specifically implemented by different linear transformations. Thereby ensuring that the query data, key data, and value data in the second query key-value pair are in the same vector space.

[0077] The second feature output by the cross-attention module in the current preset decoding layer is fused with the first fused feature to obtain a second fused feature. It can be to perform cross-layer addition processing on the feature output by the cross-attention module and the input of the cross-attention module to obtain the processed feature output by the cross-attention module. It can be understood that the feature output by the cross-attention module can be the second feature, the input of the cross-attention module can be the first fused feature, and the fusion processing of the second feature and the first fused feature can be cross-layer addition processing. Specifically, the second feature and the first fused feature are input into a second preset interaction module for fusion processing to obtain a second fused feature. Among them, the second preset interaction module can be a network module or an algorithm module provided with a network module or an algorithm module capable of performing residual connection or skip connection.

[0078] In some application scenarios, the above step S53 can be directly using the second fused feature as the feature output by the current preset decoding layer. In other application scenarios, the above step S53 can be inputting the second fused feature into a feedforward neural network for linear processing to obtain the feature output by the feedforward neural network. The feature output by the feedforward neural network and the second fused feature are input into a target preset interaction module for cross-layer addition processing to obtain the feature output by the target preset interaction module, and the feature output by the target preset interaction module is used as the feature output by the current preset decoding layer. Among them, the target preset interaction module can be a network module or an algorithm module provided with a network module or an algorithm module capable of performing residual connection or skip connection.

[0079] In some application scenarios, the above step S32 can be inputting each target fused feature into at least one linear layer respectively to obtain the text recognition result corresponding to each target fused feature. Select the text recognition result that meets the preset condition from the text recognition results corresponding to each target fused feature as the text recognition result of the image to be recognized. Among them, the preset condition can be using the character information with the highest confidence in the text recognition result corresponding to the target fused feature as the text recognition result of the image to be recognized.

[0080] Please refer to Figure 6 , Figure 6 which is the flowchart of an embodiment of the text recognition method of the present application Figure Six .

[0081] In some embodiments, the above step S32 may include the following steps: Step S61: Based on each target fused feature, determine the confidence ranking corresponding to each target fused feature, and use the target fused feature with the largest confidence ranking as the target feature. Step S62: Input the target feature into the first linear layer, the second linear layer, and the third linear layer respectively for mapping processing to obtain the type of each region in the image to be recognized output by the first linear layer, the coordinate information of each region in the image to be recognized output by the second linear layer, and the character information of each region in the image to be recognized output by the third linear layer.

[0082] The structures of the first linear layer, the second linear layer, and the third linear layer are the same, but the parameters are different. Among them, the first linear layer is used to perform scene classification on each input target fusion feature to obtain the types of regions in the image to be recognized. The type of each region in the recognized image includes a foreground type or a background type. Among them, the second linear layer is used to determine coordinate information for each input target fusion feature to obtain the coordinate information of each region in the image to be recognized. Among them, the third linear layer is used to determine character information for each input target fusion feature to obtain the coordinate information of each region in the image to be recognized.

[0083] In some other application scenarios, the above step S32 may be to input the target fusion feature into a preset linear layer to obtain the confidence level corresponding to the target fusion feature. Based on the confidence level ranking of the target fusion features, select the target fusion features that meet the preset ranking from each target fusion feature as the target features. The preset ranking may be the top m in the confidence level ranking of the target fusion features. Among them, the confidence level ranking of the target fusion features is obtained from the confidence levels corresponding to each target fusion feature. The top m may be 1. Perform at least one preset post-processing on the target features to obtain the text recognition results after each preset post-processing. Among them, when the text recognition results are at least one of the types of regions in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized, various preset post-processings match the type judgment of the regions, the determination of the coordinate information of the regions, and the determination of the character information of the regions in the text recognition results.

[0084] Please refer to Figure 7 , Figure 7 which is the flowchart of an embodiment of the text recognition method of this application Figure Seven .

[0085] In some embodiments, the above text recognition method may further include a training step for the text recognition model. The training step includes: Step S71: Obtain a sample image and the sample recognition result corresponding to the sample image. The sample recognition result is obtained through pre-annotation, and the sample recognition result includes real character information and a text-guided query vector related to the real character information. Step S72: Based on the sample image and the text-guided query vector, obtain a character-guided query vector corresponding to the sample image. Step S73: Use the sample image, the text-guided query vector, and the character-guided query vector as the input of the text recognition model to obtain the predicted recognition result corresponding to the sample image. Step S74: Based on the difference between the sample recognition result and the predicted recognition result, obtain the target loss. Step S75: Use the target loss to adjust the parameters in the text recognition model to obtain the trained text recognition model.

[0086] The sample image is of the same type as the image to be recognized above, or of a different type. The sample image contains the true character information corresponding to the text object. The specific limitations of the sample image can refer to the image to be recognized above, which will not be elaborated here. The above step S71 may include constructing features for the sample image to obtain the text-guided query vector of the sample image. The text-guided query vector of each sample image may be, for each sample image, segmenting the true character information of the sample image to obtain the index sequence of the sample image. Performing a masking operation on the index sequence of the sample image to obtain the masked index sequence corresponding to the sample image. Performing a mapping operation on the masked index sequence to obtain the text-guided query vector of the sample image.

[0087] Exemplarily, the text-guided query vector construction module mainly includes a tokenizer and a masker. The process of constructing the text-guided query vector may include: segmenting all the texts of each sample image through the sub-word tokenizer of a common language model. Specifically, for each sample image, first obtain the token index sequence corresponding to each text in the sample image, and connect different texts using the [SEP] index, so that each image corresponds to an index sequence. When training the text recognition model, pad according to the longest sequence length T in the same batch of images to ensure that the sequence lengths of a batch are the same, and use the [PAD] index as the padding flag. Performing a random masking operation on the index sequence corresponding to each text in the sequence. Specifically, randomly replace the word index of each text index sequence with the [MASK] index. To ensure the semantic information of the text, the value should not be too large. Map the masked complete sequence to the corresponding text embedding features according to the indexes therein to obtain the text-guided query vector .

[0088] During the application process of the text recognition model, the text recognition model includes an acquisition module, a preset perceptron module, a multi-layer decoder, and at least one linear layer. The acquisition module is used to acquire the input of the text recognition model. The preset perceptron module includes a first perceptron module and a second perceptron module. Among them, the input of the preset perceptron module can be the initial image features, and the output of the preset perceptron module can be the position query vectors corresponding to the initial image features. The input of the multi-layer decoder can be the initial image features and the position query vectors corresponding to the initial image features, and the output of the multi-layer decoder can be the target fusion features. The input of at least one linear layer can be the target fusion features or the sample target fusion features, and the output of at least one linear layer can be the text recognition result of the image to be recognized or the predicted recognition result of the sample image.

[0089] Please refer to Figure 8a , Figure 8aIt is a schematic diagram of a preset decoding layer in an embodiment of the text recognition method of the present application.

[0090] In some application scenarios, the multi-layer decoder in the text recognition model further includes an intra-group attention module for targets, an intra-group attention module for assistance, and an inter-group attention module. During the application process of the text recognition model, that is, in the above step S31, only the intra-group attention module for targets among the intra-group attention module for targets, the intra-group attention module for assistance, and the inter-group attention module is used as the above intra-group attention module, and only the intra-group attention module for targets participates in the processing calculations in the multi-layer decoder. During the training process of the text recognition model, the intra-group attention module for targets among the intra-group attention module for targets, the intra-group attention module for assistance, and the inter-group attention module all participate in the processing calculations in the multi-layer decoder. The input of the intra-group attention module for targets is candidate features. When the inter-group attention module participates in the processing calculations of the multi-layer decoder, the input of the inter-group attention module is the output of the intra-group attention module for targets. When the inter-group attention module does not participate in the processing calculations of the multi-layer decoder, the output of the intra-group attention module for targets is directly used as the input of the first preset interaction module. The input of the intra-group attention module for assistance is a text-guided query vector. In addition, the preset interaction module in any preset decoding layer in the multi-layer decoder can all be an Add&Norm module.

[0091] During the training process of the text recognition model, the above step S72 can be to input the sample image and the text-guided query vector into the text recognition model to obtain the features output by the multi-layer decoder in the text recognition model. It can be understood that at this time, it is the first time that the sample image and the text-guided query vector are input into the text recognition model. And the features output by the multi-layer decoder when they are first input into the text recognition model are used as the character-guided query vector corresponding to the sample image. In the case of the first input into the text recognition model, the sample image and the text-guided query vector sequentially pass through the above-mentioned acquisition module, preset perceptron module, and multi-layer decoder, and the features output by the multi-layer decoder are used as the character-guided query vector corresponding to the sample image. The above step S73 can be to obtain the sample position query vector by passing the sample initial image features of the sample image through the preset perceptron module. And the sample position query vector is used as the candidate feature. It can be understood that in the case of the first input into the text recognition model, the input of the intra-group attention module in the target group is the candidate feature, and the candidate feature is only the sample position query vector. In the case of the first input into the text recognition model, the intra-group attention module does not participate in the processing calculation of the multi-layer decoder. After determining the character-guided query vector, the determined state is that it is not the first time to input into the text recognition model. In the case of not the first time to input into the text recognition model, the intra-group attention module does not participate in the processing calculation of the multi-layer decoder. In the case of not the first time to input into the text recognition model, the intra-group attention module in the target intra-group attention module, auxiliary intra-group attention module, and inter-group attention module all participates in the processing calculation in the multi-layer decoder. In the case of not the first time to input into the text recognition model, the output of the multi-layer decoder is used as the sample target fusion feature. Based on each sample target fusion feature, the predicted recognition result of the sample image is determined. Specifically, the process of steps S61 to S62 above can be referred to determine the predicted recognition result of the sample image, which will not be elaborated here. The predicted recognition result of the sample image includes at least one of the following: the sample type of each region in the sample image, the sample coordinate information of each region in the sample image, and the sample character information of each region in the sample image.

[0092] The above step S74 can be to determine the target loss function according to the difference between the true character information in the sample recognition result and the sample character information of each region in the predicted recognition result of the sample image. It can be considered that in the case of a sufficient amount of data for the sample image, the determined target loss is used to train the parameters in the text recognition model until the scene detection model converges to obtain the trained text recognition model. The trained text recognition model is used to execute the above steps S11 to S13.

[0093] Exemplarily, a loss function is constructed, and a scene text recognition model is trained using a data set. When training the text recognition model, only the sample position query vector follows the strategy of assigning labels using the Hungarian algorithm in DETR, and the character-guided query vector calculates the loss of the corresponding task based on the pre-assigned labels. The loss functions calculated for the outputs of the sample position query vector and the character-guided query vector include: the channel dimension C of the sample type output by the classification head in the first linear layer uses the FocalLoss loss function, the sampling point coordinates P in the sample coordinate information output by the regression head in the second linear layer uses the L1 Loss, and the sample character information T output by the text head in the third linear layer uses the standard cross entropy loss function for supervision. The output TC of the text-guided query vector is supervised using the cross entropy loss function, and the corresponding label is an index sequence without masking operation.

[0094] Exemplarily, the training of the text recognition model also includes constructing a character-guided query vector. In some application scenarios, when training the text recognition model, the features of the sample position query vector first output by the multi-modal multi-layer decoder are used as the character-guided query vector. In other application scenarios, after the features of the sample position query vector first output by the multi-layer decoder are post-processed and selected by the confidence threshold, p candidate results remain from the k candidate results. The sampling points of the Bezier curves of the upper and lower edges of the predicted text form a polygon surrounding the text, and the IOU of the predicted result and the label polygon is calculated, and the p results are assigned to the corresponding labels. The text content of the p predicted results is compared with the text content of the corresponding label, and the predicted results are corrected by the minimum edit distance. The characters that have changed before and after the correction are referred to by the [UNK] symbol, and then the processed text results are input into the character segmenter to obtain the embedding vector corresponding to each character to obtain the character query vector. .

[0095] In some application scenarios, the training of the text recognition model also includes constructing a multi-modal multi-layer decoder, which consists of multiple preset decoding layers with the same structure and three task heads consisting of linear layers. , sample location query vector , text-guided query vector and character-guided query vector After inputting the output of the last decoding layer into the multimodal decoder and then inputting it into the respective task heads, the predicted text result of the sample image can be obtained. The specific steps can be referred to as follows:

[0096] The structure of each preset decoding layer is as follows Figure 8aAs shown, among them, the cross-attention module can be a structure in Transformer, not a specific structure. The first preset interaction module, the second preset interaction module, and the target preset interaction module can all be Add&Norm modules, which can perform addition operations and layer normalization operations on features. The feed-forward neural network can be expressed as FFN. The character query vector is obtained by transforming the model output prediction result. Therefore, an attention mask graph needs to be used in the calculation of the inter-group attention module, that is In the first input, the middle part of the attention mask graph is adopted to learn the sample position query The relationship between different query vectors in, and the sample position query The second time and the character-guided query vector When input together, the calculation in the first input can be omitted, and the upper half of the attention mask graph is adopted to make the character-guided query vector Can learn the sample position query vector related to its own character Information. This operation ensures that the sample position query vector follows the causal relationship and is not affected by the output result. The black squares in the mask graph represent masks and do not calculate attention.

[0097] Text-guided query vector Calculate the semantic relationship within each text in the image through the intra-group attention module. Different texts are separated by [SEP]. At least part of the attention mask graph is used in the intra-group attention module to control each text to only learn the semantic relationship between its own internal word vectors and not pay attention to other text information. The target intra-group attention module and the auxiliary intra-group attention module share parameters, so that the auxiliary intra-group attention module can simultaneously learn the feature information and semantic information of the sample position query vector within the same text. Text-guided query vector Helps the network understand the text semantics by predicting the correct word segmentation index.

[0098] After the sample position query vector, the text-guided query vector, and the character-guided query vector pass through the target intra-group attention module, the auxiliary attention module, and the inter-group attention module, the output of the first preset interaction module and the sample initial image feature Input the cross-attention module for cross-attention calculation, so that the sample position query vector And the character-guided query vector Perceive information related to the text position and each constituent character in the sample initial image feature Let the text-guided query vector In the sample initial image feature In the interaction, semantic association is performed between the text semantic information and the corresponding visual features to improve the ability of the multi-modal decoder to learn the text context semantic information and the accuracy of text recognition. The outputs of the three are passed through a feed-forward neural network FFN and a target preset interaction module Add&Norm after the output of the cross-attention module, and then output to the last preset decoding layer, respectively obtaining the sample target fusion features corresponding to the sample position query vector , the sample target fusion features corresponding to the text-guided query vector , the sample target fusion features corresponding to the character-guided query vector . The output of the last preset decoding layer , are input into the classification head, regression head, and text head composed of linear layers, respectively obtaining the class confidence corresponding to each decoded query vector , the sampling point coordinates of the upper and lower edge Bezier curves , the text content , where m is the total number of character categories in the character library plus 2. The output of the decoding layer is input into the word classification head composed of linear layers to predict the class index of each token vector , and N represents the total number of word categories in the token library.

[0099] Please refer to Figure 8b , Figure 8b which is a schematic diagram of a sample image in an embodiment of the text recognition method of the present application.

[0100] The sample image is any scene text image data in the dataset. The sample image contains text positions and character annotations. As Figure 8b shown, the sample images in the dataset are scene images containing text, and the labels of each sample image are converted into a format including the 4 control point coordinates of the Bezier curve of the text upper and lower edges in the image and character annotations. The character annotation is also the true character information included in the above sample recognition result. Specifically, each sample image includes at least character annotations and Bezier curves. In some other application scenarios, each sample image also contains the pattern elements it has. Among them, Figure 8b the points P1 to P4 in

[0101] Please refer to Figure 8c , Figure 8c which is a schematic diagram of the training module in an embodiment of the text recognition method of the present application.

[0102] As Figure 8cThe training module shown is used to train the text recognition model of the present application. Among them, the training module is used to execute the above steps S71 to S75 to obtain a trained text recognition model. The text recognition model includes a multi-scale feature extraction module, a preset feature encoder, a visual query vector construction module, a multi-layer decoder, a text-guided query vector construction module, and a character-guided query vector construction module. Among them, the sample image passes through the multi-scale feature extraction module and the preset feature encoder in sequence to obtain each initial image feature in the initial image feature set. Each initial image feature passes through the visual query vector construction module to obtain the sample position query vector Q of the sample image V The true character information passes through the text-guided query vector construction module to obtain the text-guided query vector Q related to the true character information T 。The output obtained by inputting each initial image feature, the sample position query vector Q of the sample image V and the text-guided query vector Q related to the true character information T into the multi-layer decoder is used as the input of the character-guided query vector construction module to obtain the character-guided query vector Q C 。The output obtained by inputting each initial image feature, the sample position query vector Q of the sample image V and the text-guided query vector Q related to the true character information T and the character-guided query vector Q C into the multi-layer decoder is used as the sample target fusion feature, and the sample target fusion feature is processed through the above at least one linear layer to obtain the predicted recognition result of the sample image.

[0103] Please refer to Figure 8d , Figure 8d which is a schematic diagram of the visual query vector construction module in an embodiment of the text recognition method of the present application.

[0104] Exemplarily, the sample initial image feature set includes several sample initial image features. Constructing the sample position query vector corresponding to each sample initial image feature can be formed by adding the sample initial position query vector and the sample content query vector 。Among them, each sample initial image feature outputs the probability score of each sample initial image feature through the classification linear layer Linear, and obtains the index of the top k features ranked from high to low according to the score. Each initial image feature inputs the center Bezier curve of each text instance corresponding to the L sampling point coordinates obtained by the Bezier coordinate converter composed of the first perceptron module MLP, where L refers to the unified text length. According to the top k indexes, the L sampling point coordinates of the center Bezier curve corresponding to the features ranked top k according to the score are obtained , and the sampling point coordinates The coordinate information of sampling points as initial image features. After passing the coordinate information of sampling points through cosine position encoding and inputting it into the position encoder composed of the second perceptron module, a sample initial position query vector is obtained. Adding the sample position query vector and the sample content query vector gives the sample position query vector corresponding to each sample's initial image feature. The sample content query vector is obtained by random initialization.

[0105] Please refer to Figure 8e , Figure 8e which is a schematic diagram of the text-guided query vector construction module in an embodiment of the text recognition method of this application.

[0106] Exemplarily, the text-guided query vector construction module mainly includes a tokenizer and a masker. The process of constructing the text-guided query vector may include: tokenizing all the texts of each sample image through the sub-word tokenizer of a common language model. Specifically, for each sample image, first obtain the token index sequence corresponding to each text in the sample image, and connect different texts using the [SEP] index, so that each image corresponds to an index sequence. When training the text recognition model, pad it according to the longest sequence length T in the same batch of images to ensure that the sequence lengths of a batch are the same, and use the [PAD] index as the padding flag. Perform a random masking operation on the index sequence corresponding to each text in the sequence. Specifically, randomly replace the word index of each text index sequence with the [MASK] index. To ensure the semantic information of the text, the value of a should not be too large. Map the masked complete sequence to the corresponding text embedding features according to the indices therein to obtain the text-guided query vector. .

[0107] Exemplarily, the preset text can represent the true character information in the sample image. The preset text passes through the tokenizer and the masker in sequence to obtain the text-guided query vector. The output of the preset text passing through the tokenizer is the padded token index sequence QT'. The result obtained by performing a random masking operation on the padded token index sequence is used as the text-guided query vector.

[0108] Please refer to Figure 8f , Figure 8f which is a schematic diagram of the character-guided query vector construction module in an embodiment of the text recognition method of this application.

[0109] As Figure 8fAs shown, the starting sequence includes a first sequence and a second sequence. The starting sequence can be the text prediction result obtained by passing the features first output by the multi-layer decoder for the sample position query vector and the text-guided query vector through the target linear layer. The starting sequence passes through a masker to obtain an intermediate sequence. The intermediate sequence includes the padded first sequence and the new second sequence. The starting sequence passes through a masker to pad the first sequence to obtain the padded first sequence. The padded first sequence includes the first sequence and a preset padding. The preset padding can be the [PAD] index corresponding to the above padding flag. The starting sequence passes through a masker to perform a random masking operation on the first sequence to obtain a new second sequence. The intermediate sequence passes through a tokenizer for mapping to obtain a final sequence, and the final sequence is used as the above character-guided query vector. The final sequence includes the processed first sequence and the processed second sequence. Exemplarily, the first sequence can represent "MALIBU". The second sequence can represent "27 MILFS OF SCENIC PEAUTY". The padded first sequence can be represented as "MALIBU + preset padding", that is, "MALIBU[PAD]". The new second sequence can be represented as "27 MIL[UNK]SOF SCENIC [UNK]EAUTY".

[0110] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed on each initial image feature to obtain a position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0111] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the text recognition device of the present application. The text recognition device 90 includes an acquisition module 91, an encoding module 92, and a determination module 93; the acquisition module 91 is used to obtain at least one initial image feature by using a text recognition model, and each initial image feature is obtained by performing feature extraction on the image to be recognized; the encoding module 92 is used to perform position encoding on each initial image feature to obtain a position query vector corresponding to each initial image feature; the determination module 93 is used to determine the text recognition result of the image to be recognized based on each initial image feature and each position query vector, and the text recognition result includes at least one of the following: the type of each region in the image to be recognized, the coordinate information of each region in the image to be recognized, and the character information of each region in the image to be recognized.

[0112] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed based on each initial image feature to obtain a position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0113] For the functions executed by each module, please refer to the text recognition method and will not be elaborated here.

[0114] Please refer to Figure 10 , Figure 10 FIG. is a schematic structural diagram of an embodiment of an electronic device according to the present application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is configured to execute program instructions stored in the memory 101 to implement the steps in the above text recognition method embodiment. In a specific implementation scenario, the electronic device 100 may include, but is not limited to: a multi-camera device, a microcomputer, a server. In addition, the electronic device 100 may also include mobile devices such as a laptop computer and a tablet computer, which are not limited here.

[0115] Specifically, the processor 102 is configured to control itself and the memory 101 to implement the steps in the above text recognition method embodiment. The processor 102 may also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. In addition, the processor 102 may be implemented jointly by integrated circuit chips.

[0116] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed based on each initial image feature to obtain a position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0117] Please refer to Figure 11 , Figure 11This is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 110 stores program instructions 1101 thereon. When the program instructions 1101 are executed by a processor, the steps in any of the above-described text recognition method embodiments are implemented.

[0118] In the above solution, at least one initial image feature is obtained by using a text recognition model, and then position encoding is performed based on each initial image feature to obtain a position query vector corresponding to each initial image feature. Among them, each position query vector and each initial image feature can combine feature information of different scales and different positions, thereby improving the accuracy of determining the text recognition result of the image to be recognized.

[0119] In some embodiments, the functions or modules included in the device provided by the present disclosure embodiment can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0120] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0121] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0122] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0123] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

Claims

1. A text recognition method, characterized in that, The method includes: Obtaining at least one initial image feature by using a text recognition model, where each of the initial image features is obtained by performing feature extraction on an image to be recognized; Performing position encoding on each of the initial image features to obtain a position query vector corresponding to each of the initial image features, including: for each of the initial image features, performing the following steps: inputting the initial image feature into a first perceptron module to obtain sampling point coordinate information of the initial image feature output by the first perceptron module, where the sampling point coordinate information represents position information of a preset number of sampling points of a Bessel curve related to the image to be recognized in the image to be recognized; using the preset-encoded sampling point coordinate information as the input of a second perceptron module to obtain a position query vector corresponding to the initial image feature output by the second perceptron module; Based on each of the initial image features and each of the position query vectors, determining a text recognition result of the image to be recognized, where the text recognition result includes at least one of the following: types of regions in the image to be recognized, coordinate information of regions in the image to be recognized, and character information of regions in the image to be recognized.

2. The method according to claim 1, characterized in that, The determining the text recognition result of the image to be recognized based on each of the initial image features and each of the position query vectors includes: Inputting each of the initial image features and each of the position query vectors into a multi-layer decoder for feature fusion to obtain a target fusion feature corresponding to each of the initial image features; Based on each of the target fusion features, determining the text recognition result of the image to be recognized.

3. The method according to claim 2, wherein The multi-layer decoder includes a plurality of preset decoding layers arranged in cascade, where structures of the preset decoding layers are the same but parameters are different, and using the feature output by the last preset decoding layer as the feature output by the multi-layer decoder, where the feature output by the multi-layer decoder is a target fusion feature corresponding to each of the initial image features, and each of the preset decoding layers includes an intra-group attention module; The inputting each of the initial image features and each of the position query vectors into a multi-layer decoder for feature fusion to obtain a target fusion feature corresponding to each of the initial image features includes: Using a first query-key-value pair as the input of the intra-group attention module in the current preset decoding layer to obtain a first feature output by the intra-group attention module in the current preset decoding layer, where the query data, key data, and value data in the first query-key-value pair are respectively obtained from each of the position query vectors or from the feature output by the previous preset decoding layer, and the current preset decoding layer is one of the plurality of preset decoding layers; Fusing the first feature output by the intra-group attention module in the current preset decoding layer with the input of the intra-group attention module in the current preset decoding layer to obtain a first fusion feature; Performing advanced fusion processing on the first fusion feature and each of the initial image features to obtain the feature output by the current preset decoding layer.

4. The method according to claim 3, characterized in that, Each of the preset decoding layers includes a cross-attention module. The advanced fusion processing of the first fusion feature and each of the initial image features to obtain the feature output by the current preset decoding layer includes: Using the second query-key-value pair as the input of the cross-attention module in the current preset decoding layer to obtain the second feature output by the cross-attention module in the current preset decoding layer. The query data in the second query-key-value pair is obtained from the first fusion feature, and the key data and value data in the second query-key-value pair are obtained from each of the initial image features respectively; Fusing the second feature output by the cross-attention module in the current preset decoding layer with the first fusion feature to obtain a second fusion feature; Based on the second fusion feature, obtaining the feature output by the current preset decoding layer.

5. The method according to claim 2, wherein The determining of the text recognition result of the image to be recognized based on each of the target fusion features includes: Based on each of the target fusion features, determining the confidence ranking corresponding to each of the target fusion features, and using the target fusion feature with the largest confidence ranking as the target feature; Inputting the target feature into the first linear layer, the second linear layer, and the third linear layer respectively for mapping processing to obtain the type of each region in the image to be recognized output by the first linear layer, the coordinate information of each region in the image to be recognized output by the second linear layer, and the character information of each region in the image to be recognized output by the third linear layer.

6. The method according to any one of claims 1 to 4, characterized in that, The obtaining of at least one initial image feature by using the text recognition model includes: Using the text recognition model to obtain the image to be recognized; Performing multi-scale feature extraction on the image to be recognized to obtain a multi-scale feature set corresponding to the image to be recognized. The multi-scale feature set contains at least two scale features and the scales of each of the scale features are different; Fusing each of the scale features in the multi-scale feature set to obtain the at least one initial image feature.

7. The method according to any one of claims 1 to 4, characterized in that, It further includes a training step for the text recognition model. The training step includes: Obtaining a sample image and the sample recognition result corresponding to the sample image. The sample recognition result is obtained through pre-annotation and includes real character information and a text-guided query vector related to the real character information; Based on the sample image and the text-guided query vector, obtaining a character-guided query vector corresponding to the sample image; Using the sample image, the text-guided query vector, and the character-guided query vector as the input of the text recognition model to obtain the predicted recognition result corresponding to the sample image; Based on the difference between the sample recognition result and the predicted recognition result, obtaining a target loss; Using the target loss to adjust the parameters in the text recognition model to obtain a trained text recognition model.

8. An electronic device, characterized in that, It includes: A memory and a processor. Among them, the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1-7.

9. A computer-readable storage medium having program instructions stored thereon, characterized in that, When executed by a processor, the program instructions are used to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Scene text positioning method and device, medium and product

    CN117935270A

  • Scene text recognition method based on improved Transform network

    CN118570788A