Text recognition method and storage medium

By performing text segmentation and feature mask sampling on the image, the recognition accuracy and efficiency of optical character recognition technology in different characters and environments is solved, and an efficient text recognition method is realized.

CN119785360BActive Publication Date: 2025-07-11ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510263934.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-11
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing optical character recognition technology has low recognition accuracy and efficiency when facing the influence of different character styles, fonts and external environments, especially the computing and resource consumption of large language models based on Transformer architecture.

Method used

By performing text segmentation processing on the image to be recognized, an image feature mask is generated, sampling is performed, and after reducing the feature amount, inputting the pre-trained attention network for feature extraction, the target text is determined in combination with a large language model.

Benefits of technology

It reduces the data processing volume of neural networks, improves the accuracy and efficiency of text recognition, and adapts to the text recognition needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785360B_ABST
    Figure CN119785360B_ABST
Patent Text Reader

Abstract

The present application discloses a text recognition method and a storage medium. The method includes: performing text segmentation processing on the image to be recognized to obtain a segmented image containing text regions; performing binary marking processing on the segmented image according to the pixel information of the text regions to obtain an image feature mask; performing sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampled feature; inputting the sampled feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain a target image feature; and determining a corresponding target text according to the target image feature. The above solution can improve the accuracy of text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a text recognition method and a storage medium. Background Art

[0002] Optical Character Recognition (OCR) has a large number of usage scenarios in real life. However, due to different styles and fonts of characters, as well as the influence of various external environments, it is easy for the text recognition model to have problems such as performance degradation or misrecognition.

[0003] Currently, the technology of large language models (LLMs) based on the Transformer architecture has achieved rapid development. Its core principle is to use the self-attention mechanism to capture the dependencies in the input sequence, so as to process sequence-to-sequence tasks, such as machine translation, text generation, and question-and-answer systems. On this basis, in the visual field, with the same Vision Transformer method, by virtue of its global attention mechanism, the generality and robustness of various methods in the visual field have been greatly improved, such as recognizing text information in images.

[0004] However, the computational complexity of the self-attention mechanism is positively correlated with the number of tokens (the basic unit of the input sequence). Too many token numbers will not only make the model training process difficult, but also lead to a more time-consuming and resource-consuming training process and inference process, and reduce the efficiency of the attention mechanism. Summary of the Invention

[0005] This application provides at least one text recognition method, device, equipment, and computer-readable storage medium.

[0006] In the first aspect of this application, a text recognition method is provided, including: performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region; performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask; performing sampling processing on the to-be-recognized features of the image to be recognized according to the image feature mask to obtain sampling features; inputting the sampling features and the to-be-recognized features into a pre-trained attention network for feature extraction processing to obtain target image features; and determining corresponding target text according to the target image features.

[0007] In one embodiment, the image feature mask includes a foreground region and a background region. Sampling the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampling feature includes: determining the number of samplings according to the number of foreground regions; performing discrete sampling processing on the feature to be recognized of the image to be recognized according to the number of samplings to obtain the sampling feature.

[0008] In one embodiment, before performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region, the method further includes: performing size adjustment processing on the received original image to obtain a to-be-processed image matching a preset size; performing segmentation embedding processing on the to-be-processed image to obtain the image to be recognized and the feature to be recognized of the image to be recognized.

[0009] In one embodiment, performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region includes: inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for segmentation processing and font standardization processing to obtain the segmented image; the font types of each text region in the segmented image are the same.

[0010] In one embodiment, inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for segmentation processing and font standardization processing to obtain the segmented image includes: inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for downsampling and feature extraction processing to obtain a downsampled feature of the feature to be recognized; performing upsampling processing on the downsampled feature through the image segmentation network to obtain a segmented image including the text region.

[0011] In one embodiment, performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask includes: performing partitioning processing on the segmented image to obtain at least one image region; respectively determining a binary mark corresponding to each image region according to the pixel category of each pixel point in each image region to obtain the image feature mask.

[0012] In one embodiment, the segmented image includes the text region and the non-text region. The pixel points in the text region are foreground pixels, and the pixel points in the non-text region are background pixels. Determining the binary markers corresponding to each image region respectively according to whether each pixel point in each image region is in the text region to obtain the image feature mask includes: obtaining the foreground quantity ratio of the foreground pixels and the background quantity ratio of the background pixels in each image region; performing first marking processing on the image regions where the foreground quantity ratio is greater than the background quantity ratio to obtain a first marking value; performing second marking processing on the image regions where the foreground quantity ratio is less than or equal to the background quantity ratio to obtain a second marking value; the first marking value and the second marking value are different; determining the image feature mask according to the first marking value and the second marking value.

[0013] In one embodiment, inputting the sampled feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain the target image feature includes: performing linear transformation processing on the sampled feature to obtain a query vector; performing linear transformation processing on the feature to be recognized to obtain a key vector and a value vector; performing feature extraction processing on the query vector, the key vector, and the value vector through the attention network to obtain the target image feature.

[0014] In one embodiment, determining the corresponding target text according to the target image feature includes: performing feature extraction processing on the received prompt text to obtain prompt text features; performing feature splicing processing on the prompt text features and the target image feature to obtain spliced features; inputting the spliced features into a pre-trained large language model to obtain the target text output by the large language model.

[0015] The second aspect of the present application provides a text recognition device, including: a text prediction module for performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region; a mask generation module for performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask; a sampling module for sampling the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampled feature; a feature extraction module for inputting the sampled feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain a target image feature; a text determination module for determining the corresponding target text according to the target image feature.

[0016] The third aspect of the present application provides an electronic device, including a memory and a processor, and the processor is used to execute the program instructions stored in the memory to implement the above text recognition method.

[0017] In the fourth aspect of the present application, a computer-readable storage medium is provided, on which program instructions are stored, and when the program instructions are executed by a processor, the above-mentioned text recognition method is implemented.

[0018] In the above solution, by performing text segmentation processing on the image to be recognized, a segmented image containing text regions can be obtained; according to the pixel information of the text regions, binary marking processing is performed on the segmented image to obtain an image feature mask with foreground and background separated; according to the image feature mask, sampling processing is performed on the features to be recognized of the image to be recognized, and then the sampling features of the text regions in the image to be recognized can be sampled; the sampling features and the features to be recognized are input into a pre-trained attention network for feature extraction processing, and the target image features output by the attention network are obtained, and then the corresponding target text can be determined according to the target image features. Thus, the data processing amount of the neural network in the subsequent process can be reduced by the method of image sampling processing, and the accuracy of text recognition can also be improved.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Brief Description of the Drawings

[0020] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to explain the technical solutions of the present application.

[0021] Figure 1 is a schematic flowchart of an exemplary embodiment of the text recognition method of the present application;

[0022] Figure 2 is a schematic diagram of text font de-stylization in an exemplary text recognition method of the present application;

[0023] Figure 3 is a schematic diagram of an image semantic block in an exemplary text recognition method of the present application;

[0024] Figure 4 is a schematic diagram of an exemplary text recognition model in the text recognition method of the present application;

[0025] Figure 5 is a block diagram of a text recognition device shown in an exemplary embodiment of the present application;

[0026] Figure 6 is a schematic structural diagram of an embodiment of an electronic device of the present application;

[0027] Figure 7 is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. Detailed Description of the Embodiments

[0028] The solution of the embodiment of the present application will be described in detail below in conjunction with the accompanying drawings of the specification.

[0029] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0030] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the preceding and following associated objects. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article means any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0031] For ease of understanding, an exemplary description of one of the applicable scenarios of the present application is now given.

[0032] Optical Character Recognition technology (OCR) has a large number of usage scenarios in real life. However, due to different character styles, fonts, etc., and the influence of various external environments, it is easy for the text recognition model to experience performance degradation or misrecognition. In recent years, the technology of large language models (LLMs) based on the Transformer architecture has achieved rapid development. On this basis, in the field of computer vision, with the same Vision Transformer method, by virtue of its global attention mechanism, the generality and robustness of various methods in the visual field have been greatly improved. Therefore, the large language model based on the Transformer architecture can be used in text recognition methods.

[0033] However, the computational complexity of the self-attention mechanism is proportional to the square of the number of tokens. An excessive number of tokens can make the model training and model inference processes more difficult and incur higher costs. Here, a token refers to the smallest semantic unit representing a word in a language model. Before the prompt text is sent to the neural network, the Tokenizer breaks down long texts such as compound words, sentences, paragraphs, and articles into token units, which are then transformed into vector representations through Embedding and finally input into the neural network. It can be understood that when recognizing text in an image, the image needs to be decomposed into image tokens, and the feature vectors of the image tokens are input into the neural network to determine the text information therein.

[0034] In daily scenarios, it can be observed that the number of words in different images shows randomness in distribution: for example, there may be a large number of advertising slogans in mall images, while text slogans are rarely seen in remote outdoor scenarios. With the continuous upgrade of image acquisition devices, their imaging resolution is getting better and better, and the information contained in the images tends to be more diverse and complex.

[0035] Therefore, it is not very practical to use the same text recognition method to recognize text in images of different scenarios. How to adaptively recognize the text in an image according to the information already contained in the image is of certain significance for text recognition in natural scenarios.

[0036] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an exemplary embodiment of the text recognition method of this application. Specifically, it may include the following steps:

[0037] Step S110: Perform text segmentation processing on the image to be recognized to obtain a segmented image containing the text region.

[0038] Performing text prediction processing on the image to be recognized means classifying each pixel point in the image to be recognized and predicting whether it represents the text region in the image to be recognized based on its category.

[0039] Exemplarily, in this application, the segmentation network TextSegNet can be used to perform text segmentation processing on the image to be recognized to obtain the segmented image . Specifically, the image to be recognized can be input as the input image of the segmentation network, and then the Seg-Encoder in the segmentation network performs downsampling and feature extraction on the image to be recognized to obtain the deep features of the image to be recognized ; Then, the deep features extracted are upsampled by the Seg-Decoder in the segmentation network, and the segmented image can be obtained. .

[0040] The core principle of SegNet is to achieve semantic segmentation of images through deep learning methods. Different from traditional pixel-level classification methods, SegNet can not only classify each pixel in the image, but also assign each pixel to a predefined category, thus achieving refined segmentation of the image, and predicting the text regions and non-text regions in the image to be recognized. Among them, the encoder is responsible for extracting deep features from the original image, and the decoder uses these features to generate the final semantic segmentation map. The Seg-Encoder in the segmentation network can use a stacked Convnet to downsample and extract features from the image to be recognized, and the decoder gradually restores the resolution through upsampling and deconvolution operations, and reconstructs the semantic segmentation of the input image according to the features extracted by the encoder. Details are not elaborated here.

[0041] It should be noted that when performing the above method, it can also be to input the features to be recognized of the image to be recognized into the segmentation network for downsampling, feature extraction, upsampling, etc.

[0042] Optionally, it can also be to use a method based on pixel segmentation in image processing (such as according to pixel information, edge information, connected component information, etc.) of the image to be recognized to determine the text regions and non-text regions in the image to be recognized, and obtain the segmented image. Details are not limited here.

[0043] Step S120, perform binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask.

[0044] Combined with the foregoing steps, by performing segmentation processing on the image to be recognized, the segmented image can be obtained , and according to the segmented image the foreground and background in the segmented image can be determined according to the pixel information (such as the category of each pixel) of each pixel in it, that is, it is equivalent to determining the text regions and non-text regions in the segmented image.

[0045] Among them, the binary marking processing in this application can be to mark the segmented image with 0 and 1, or to mark the segmented image with two other different numerical values. Details are not limited here.

[0046] Specifically, it can be marking the pixel points in the foreground of the segmented image as 1 and the pixel points in the background as 0; it can also be performing a partitioning process on the segmented image, marking the area with more foreground pixels as 1 and the area with more background pixels as 0, which is not limited here. Thus, the text area and the non-text area can be segmented and represented, achieving the determination of the image feature mask based on the pixel information in the segmented image in the .

[0047] Step S130: Sampling the to-be-recognized features of the to-be-recognized image according to the image feature mask to obtain the sampled features

[0048] Combined with the foregoing steps, after the foregoing steps of processing, the image feature mask can represent the discrimination result of the text area and the non-text area of the to-be-recognized image

[0049] Therefore, in this step, the image feature mask can be used as the weight for the sampling process, and the to-be-recognized features of the to-be-recognized image are sampled accordingly to obtain the sampled features .

[0050] It can be understood that after the sampling process of the to-be-recognized image using the image feature mask in this step, the obtained sampled features are significantly more concise than the to-be-recognized image in terms of the information contained in the image

[0051] Step S140: Input the sampled features and the to-be-recognized features into a pre-trained attention network for feature extraction processing to obtain the target image features

[0052] The attention mechanism can help the neural network model better focus on the key information in the input data. For example, in the image recognition scenario, the attention mechanism can help the neural network model better focus on the key features in the image, thereby improving the recognition accuracy and enabling more accurate text to be obtained subsequently

[0053] Combined with the foregoing steps, taking the sampled features obtained by sampling as the query subject and the to-be-recognized features of the to-be-recognized image as the queried object and inputting them into a pre-trained attention network for feature extraction processing, the target image features output by the attention network can be obtained .

[0054] Step S150: Determine the corresponding target text according to the target image features

[0055] The above steps are combined for illustration. After obtaining the target image features, the target image features can be input into a trained large language model to obtain the target text output by the large language model. Alternatively, similarity calculation is performed based on the target image features and the base image features in the preset database, and the target base feature most similar to the target image features among the base image features is determined through feature similarity. Each base image feature can have corresponding text information of words, phrases or sentences, so that the corresponding target text can be found and output according to the target image features.

[0056] It can be seen that in this application, through text segmentation processing on the image to be recognized, a segmented image containing a text area after segmentation can be obtained; binary marking processing is performed on the segmented image according to the pixel information of the text area to obtain an image feature mask with foreground and background separated; sampling processing is performed on the feature to be recognized of the image to be recognized according to the image feature mask, and the sampling feature of the text area in the image to be recognized can be sampled; the sampling feature and the feature to be recognized are input into a pre-trained attention network for feature extraction processing, and the target image features output by the attention network are obtained, and then the corresponding target text can be determined according to the target image features. In this way, the data processing amount of the neural network in the subsequent process can be reduced through the method of image sampling processing, and the accuracy of text recognition can also be improved.

[0057] Based on the above embodiments, the embodiment of this application illustrates the step of performing sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask to obtain the sampling feature. Among them, the image feature mask includes a foreground area and a background area. Specifically, the method of this embodiment includes the following steps:

[0058] Determine the number of sampling times according to the number of foreground areas; perform discrete sampling processing on the feature to be recognized of the image to be recognized according to the number of sampling times to obtain the sampling feature.

[0059] Combined with the foregoing embodiments for illustration, in the process of obtaining the image feature mask the foreground area and the background area therein are marked by binary marking processing. Among them, the foreground area is equivalent to the text area determined by the segmentation network, that is, the area to be sampled in the image to be recognized.

[0060] Therefore, the number of sampling times can be determined according to the number of foreground areas in the image feature mask (for example, determining the number of foreground areas as the number of sampling times, or determining a value less than or greater than the number of foreground areas as the number of sampling times), that is, it is equivalent to adaptively determining the number of sampling times according to the number of targets included in the image feature mask. Among them, in the case of binary marking processing with 0 and 1, the number of targets is equivalent to (the sum of the marking values of each foreground area and background area in the image feature mask).

[0061] Exemplarily, discrete sampling processing is performed on the feature to be recognized according to the number of samplings, which may be to utilize an image feature mask in matrix form for the feature to be recognized of the image to be recognized to perform discrete sampling, or other sampling methods with mathematical differentiability, which are not limited herein. Then, the sampled feature after adaptive sampling can be obtained . The mathematical expression of which can be:

[0062]

[0063] wherein represents a multinomial distribution, and uses as the weight for multinomial distribution sampling to sample the target of the potential pre-text region in the image to be recognized. R refers to the set of real numbers represents the feature dimension after mapping of each image token. m refers to the number of partitions when partitioning the segmented image to determine the image feature mask When performing non-overlapping partitioning processing on a text region of size H*W using a window with a preset pixel size, then partitions can be obtained

[0064] Based on the above embodiments, the embodiments of the present application illustrate the steps before performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region. Specifically, the method of this embodiment includes the following steps:

[0065] Adjust the size of the received original image to obtain a to-be-processed image matching the preset size; perform segmentation embedding processing on the to-be-processed image to obtain the image to be recognized and the feature to be recognized of the image to be recognized

[0066] Combined with the foregoing embodiments for illustration, the image to be recognized in the present application may be an image after image processing or an image without image processing. The process of performing text prediction processing on the image to be recognized may be to input the image to be recognized into a segmentation network for text prediction processing; or to input the feature to be recognized of the image to be recognized into the segmentation network for text prediction processing

[0067] Taking the method of inputting the feature to be recognized into the segmentation network for text prediction processing as an example for illustration, it is necessary to obtain the feature to be recognized before performing text prediction processing on the image to be recognized

[0068] Exemplarily, the image to be recognized before image processing is called the original image. The original image is subjected to size adjustment processing to obtain a to-be-processed image I that matches the preset size (H*W), which can be denoted as , for subsequent image processing. Using window to perform segmentation and embedding processing (Patch Embedding) on the to-be-processed image I. The principle is to first segment the to-be-processed image I using the above window to obtain the image to be recognized, and then embed the image to be recognized into the vector space. The encoded image tokens are denoted as , then the to-be-recognized features of the image to be recognized are obtained. Among them, represents the number of tokens, represents the feature dimension after mapping for each image token, which can be set to 512 according to experience.

[0069] Therefore, the subsequent process can be to input into the segmentation network TextSegNet. Through the Seg-Encoder of the segmentation network TextSegNet (a stacked Convnet can be used) to further downsample and extract features from , a feature map of size is obtained , and then the Seg-Decoder upsamples the feature map to obtain the predicted segmentation image , whose size is . In this embodiment, the segmentation network TextSegNet is mainly used to perform binary classification on the image I, and the category of each pixel point in the obtained segmentation map can represent whether it is a text area.

[0070] Based on the above embodiment, the embodiment of the present application exemplarily illustrates the steps of performing text segmentation processing on the image to be recognized to obtain a segmentation image including text areas. Specifically, the method of this embodiment includes the following steps:

[0071] Input the to-be-recognized features of the image to be recognized into a pre-trained image segmentation network for segmentation processing and font normalization processing to obtain a segmentation image; the font types of each text area in the segmentation image are the same.

[0072] Among them, the description of the segmentation processing can refer to the description of the previous and subsequent embodiments, which will not be elaborated here. Font normalization processing means that there may be texts of different font types (such as regular script, Song typeface, etc.) in the image to be recognized.

[0073] It should be noted that considering the diversity and complexity of text fonts in real-world scenarios, although continuously expanding the training data is a conventional and effective method (which can also be selectively implemented in this application), the continuous collection, annotation, and cleaning of data consume huge resources. Therefore, during the training process of the segmentation network of this application, instead of directly learning to generate various text fonts corresponding to the input image, it can learn the binary map after font standardization. As Figure 2 shown, Figure 2 is an exemplary schematic diagram of text font de-stylization in the text recognition method of this application. For texts of various fonts in the image, they can all be processed to a single font style type through style standardization.

[0074] Thus, when performing recognition and detection on texts of different fonts, it is possible to avoid the font style (type) from affecting the final recognition and detection results. Therefore, the image segmentation network of this application can standardize the font types of each text region in the image to be detected, making the font types of each text region the same. Specifically, reference data on font styles can be introduced during the training of the image segmentation network for model training to achieve the effect of text de-stylization, which will not be elaborated here.

[0075] Exemplarily, in the training stage, through supervised training, the parameters in the Patch Embedding processing module can be adjusted to enable the neural network to further focus on the character feature information after font de-stylization. Specifically, it can be to and calculate the binary cross-entropy loss. Considering that the difficulty of classifying each pixel point is different, when the predicted value approaches 0.5, the loss weight should be increased. From the definition of information entropy, for numbers between 0 and 1, when the value is 0.5, its entropy value is the largest. Therefore, using the information entropy of each pixel point as the weight of the binary cross-entropy loss, its mathematical expression can be:

[0076]

[0077] where represents the predicted value corresponding to the pixel position in, ranging from 0 to 1; represents the label value corresponding to the pixel position in, taking values of 0 or 1.

[0078] The main purpose of text font de-stylization is to convert texts of various styles into a certain specific standard style. Therefore, a style loss can be added as an auxiliary constraint, and its mathematical expression can be:

[0079]

[0080] Among them It represents the common style loss calculation method, which will not be elaborated here.

[0081] TextSegNet can adapt the structure based on the existing image segmentation network, directly load the corresponding trained model weights, and freeze the model parameters of this part during the training process to reduce the computing resources required for training.

[0082] By the above method, an auxiliary constraint for font de-stylization is proposed, and the text in the binary label map in the segmentation network is font-normalized. The shallow features of the network are supervised by the style loss and the weighted segmentation loss, so that the shallow stage of the network can extract the potential representative features of the characters and reduce the misrecognition caused by font styles. The segmentation network, as an auxiliary module of the text recognition method in this application, does not require the obtained binary segmentation map to be accurate enough, which can also simplify the training difficulty and further improve the robustness of the text recognition network.

[0083] Based on the above embodiments, the embodiments of this application will describe the steps of inputting the to-be-recognized features of the to-be-recognized image into a pre-trained image segmentation network for segmentation processing and font normalization processing to obtain a segmented image. Specifically, the method of this embodiment includes the following steps:

[0084] Input the to-be-recognized features of the to-be-recognized image into a pre-trained image segmentation network for downsampling and feature extraction processing to obtain the downsampled features of the to-be-recognized features; perform upsampling processing on the downsampled features through the image segmentation network to obtain a segmented image including the text region.

[0085] For the specific processing process of this embodiment, reference can be made to the description in the foregoing embodiments, which will not be elaborated here.

[0086] Based on the above embodiments, the embodiments of this application will describe the steps of performing binary labeling processing on the segmented image according to the pixel information of the text region to obtain an image feature mask. Specifically, the method of this embodiment includes the following steps:

[0087] Perform partitioning processing on the segmented image to obtain at least one image region; determine the binary label corresponding to each image region respectively according to whether each pixel point in each image region is in the text region to obtain an image feature mask.

[0088] Combined with the foregoing embodiments, after obtaining the segmented image Then it can be used The window size is used to perform non-overlapping partition processing on the segmented image to obtain at least one image region. The size of the segmented image is H*W, so the number of image regions obtained by partitioning can be The predicted classification results of each pixel in each area are counted in turn, and the window corresponding to each image area is binary labeled according to the pixel category of each pixel in each image area (whether it is in the text area) to obtain the image feature mask.

[0089] For example, according to a preset window size, the segmented image is subjected to sliding window statistical processing (determining the type of pixel points in each window of the segmented image), and each window area is binary-labeled according to the statistical results to obtain an image feature mask.

[0090] On the basis of the above embodiments, the embodiment of the present application describes the steps of determining the binary label corresponding to each image area according to whether each pixel point in each image area is in the text area, and obtaining the image feature mask. The segmented image includes a text area and a non-text area, the pixel points in the text area are foreground pixels, and the pixel points in the non-text area are background pixels. Specifically, the method of this embodiment includes the following steps:

[0091] Obtain the foreground number ratio of foreground pixels and the background number ratio of background pixels in each image area; perform a first marking process on the image area where the foreground number ratio is greater than the background number ratio to obtain a first marking value; perform a second marking process on the image area where the foreground number ratio is less than or equal to the background number ratio to obtain a second marking value; the first marking value and the second marking value are different; determine the image feature mask according to the first marking value and the second marking value.

[0092] In conjunction with the above-mentioned embodiments, the present application performs text segmentation on the image to be recognized, and can obtain the type corresponding to each pixel in the segmented image, that is, characterize whether each pixel is a foreground pixel or a background pixel (equivalent to a text area pixel or a non-text area pixel). According to the binary marking method provided in the above-mentioned embodiments, the first marking value can be 1, and the second marking value can be 0.

[0093] Specifically, when binary labeling is performed on each image region, the labeling method can be determined by referring to the type of pixels in each image region. For example, the predicted classification results of each pixel in each region are counted respectively, and the ratio of the number of foreground pixels to the number of background pixels in each region is calculated. For example:

[0094] 1. Calculate the ratio of the number of foreground pixels in each region to the total number of pixels in that region to obtain the foreground ratio. If the foreground ratio of the region is greater than 50% or greater than other preset ratio thresholds, the region can be marked with a first mark value; otherwise, the region is marked with a second mark value.

[0095] 2. Calculate the ratio of the number of foreground pixels in each region to the total number of pixels in that region to obtain the foreground ratio. Similarly, calculate the ratio of the number of background pixels in each region to the total number of pixels in that region to obtain the background ratio. Mark the regions where the foreground ratio is greater than the background ratio with a first mark value; mark the regions where the foreground ratio is less than or equal to the background ratio with a second mark value.

[0096] Segment the image through the above method Mark the regions with a higher foreground ratio in the image as 1, otherwise mark them as 0, thereby obtaining a binary mark matrix, that is, an image feature mask .

[0097] Based on the above embodiments, the embodiments of the present application illustrate the steps of inputting the sampling feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain the target image feature. Specifically, the method of this embodiment includes the following steps:

[0098] Perform a linear transformation on the sampling feature to obtain a query vector; perform a linear transformation on the feature to be recognized to obtain a key vector and a value vector; perform feature extraction processing on the query vector, key vector, and value vector through an attention network to obtain the target image feature.

[0099] Combined with the foregoing embodiments, in the process of performing feature extraction processing on the sampling feature and the feature to be recognized based on the attention network, it is necessary to perform a linear transformation on the sampling feature to obtain a query vector Query (Q), and perform a linear transformation on the feature to be recognized to obtain a key vector Key (K) and a value vector Value (V). Then, based on the Q, K, V vectors and the cross-attention mechanism, perform processing to obtain the target image feature , and its mathematical expression can be:

[0100]

[0101] Among them, , and Represents a learnable transformation matrix. Then, the obtained image features are repeatedly subjected to Group Query Attention calculation N - 1 times (N can be preset according to daily experience values and is not limited here), and the finally output image features of r can be obtained. .

[0102] Based on the above embodiments, the embodiments of the present application illustrate the steps of determining the corresponding target text according to the target image features. Specifically, the method of this embodiment includes the following steps:

[0103] According to the received prompt text, perform feature extraction processing on the prompt text to obtain prompt text features; perform feature splicing processing on the prompt text features and the target image features to obtain spliced features; input the spliced features into a pre-trained large language model to obtain the target text output by the large language model.

[0104] The prompt text refers to the text information used to prompt the large language model (LLM), which can be used to indicate or guide the large language model to generate specific outputs. In order to enable the large language model to accurately output the required data, the obtained target image features and the prompt text can be input into the large language model together, so that the large language model can obtain information, generate information and output from the target image features according to the instructions of the received prompt text to obtain the target text.

[0105] Exemplarily, for example, the text prompt can be "Please output all the texts in the image in sequence according to the semantic block information, and use semicolons as separators for texts in different semantic blocks". Among them, the semantic blocks can be a certain advertisement slogan or slogan in the corresponding image, etc. There are often intuitive differences in font, size, color and position between different semantic blocks. As Figure 3 shown, Figure 3 is an exemplary image semantic block schematic diagram in the text recognition method of the present application. AB and CD belong to the same semantic block, EFG belongs to the same semantic block, HI and JK belong to the same semantic block, and LMN belongs to the same semantic block. Therefore, after the large language model analyzes and processes according to the above text prompt and the image semantic blocks, the output target text can be "ABCD; EFG; HIJK; LMN".

[0106] Optionally, in addition to using the prompt text and the large language model to generate the target text, it can also be to directly generate the corresponding target text according to the target image features by using other text generation models or by means of feature similarity search, which is not limited here.

[0107] Specifically, the text prompt can be subjected to feature extraction processing by a text feature extractor to obtain the prompt text features denoted as , and the target image features After splicing, the spliced features are obtained and used as the input of the LLM, and then the target text output by the LLM is obtained.

[0108] It should be noted that reference can be made to Figure 4 as shown Figure 4 which is a schematic diagram of an exemplary text recognition model in the text recognition method of this application. When implementing the text recognition method of this application, the application process of this model is mainly as follows: The original image is scaled to the size, and image tokens are obtained through the Patch Embedding operation of the segmentation embedding module, and then are respectively sent into the lightweight segmentation network TextSegNet and the image encoder Vision Transformer Encoder to obtain the image feature mask and the feature to be recognized. Among them, the image encoder may include a cross-attention mechanism module (CrossAttention) and a group query attention module (Group Query Attention). Adaptive sampling processing is performed by the adaptive sampling module to determine the filtered encoded features (which can be denoted as dynamic image tokens). Finally, the encoded features of the prompt text (which can be denoted as text tokens) are combined and sent into the large language model LLM to generate corresponding prediction results.

[0109] In the model training process of this application, when the spliced features are input into the LLM for training, after the forward calculation of the LLM, the loss can be calculated between the result predicted by the large language model and the data annotation, denoted as . Then the mathematical expression of the overall loss of the text recognition network of this application can be:

[0110]

[0111] where , and are the weight parameters of the loss, which can be set to , , according to empirical values, and there is no limitation here. Among them, since the segmentation network is an auxiliary module of the text recognition model of this application, it is not required that the segmentation map output by the segmentation network is accurate enough. Therefore, less weight can be used for the loss of the segmentation network part. In the training stage, the parameters of the trainable part in the model are continuously updated by minimizing the total loss value until the model converges, and the trained neural network is obtained and used to execute the above embodiments of the text recognition process.

[0112] Based on the above embodiments, this embodiment can also provide a method for perceiving environmental information in the image to be recognized (or the initial image). Exemplarily, after obtaining the image feature mask in the foregoing embodiment, an inversion operation can be performed on the image feature mask. For example, the value marked as 1 in the image feature mask is inverted to 0, and the value marked as 0 in the image feature mask is inverted to 1 to obtain the background image mask. According to the description of the foregoing embodiment, the background image mask obtained by inverting the image feature mask can shield the information of the text area in the image. Shielding the image to be recognized with the background image mask can also refer to the foregoing sampling method by analogy to obtain the background image. The background image is input into a pre-trained environmental perception network, and the environmental feature information of the background image is extracted by the environmental perception network. Based on the environmental feature information and the target image features, the corresponding target text is determined (or based on the environmental feature information, the target image features, and the prompt text features, the corresponding target text is determined). For specific reference, please refer to the description of the foregoing embodiment, which will not be elaborated here.

[0113] Thus, in the scenario where environmental feature information is introduced, the large language model can fully interpret the image to be recognized, and the target text output by the large language model can contain richer text information. It should be noted that the large language model can not only output the environmental information in text form, such as a text description of the current image environment, or convert the environmental information into a text description form and combine it with the recognized text information for output. It can also deeply understand the image to be recognized with the help of environmental information. For example, the text to be recognized (ABCD, EFG) is on a billboard, and the text to be recognized (HIJK is on a notice board), and the text to be recognized (LMN) in the image to be recognized is on a banner. Sampling the text area through the image feature mask can accurately obtain its text information (such as ABCD, EFG, HIJK, LMN), and sampling the non-text area through the background image mask can obtain its environmental information (such as billboard, notice board, banner, etc.). Therefore, when the large language model receives the prompt text "Please output the text on the banner", it can accurately output "LMN". Another example is that when the large language model receives the prompt text "Please output the text on the banner and separate them with semicolons", it can accurately output "ABCD;EFG". In addition, it can also identify whether the text to be recognized in the current scene is indoors and / or outdoors, and more text locations according to the environmental feature information, which is not limited here.

[0114] Based on the above embodiments, this embodiment can further provide a method for determining a text region, including: obtaining the contrast information of the image to be recognized, and in response to the contrast information being higher than a preset contrast threshold, optionally using a preset edge detection algorithm or a connected component detection algorithm to determine the text region in the image to be recognized; in response to the contrast information being lower than or equal to the contrast threshold, optionally using the method of using a segmentation network in the foregoing embodiments to determine the text region in the image to be recognized.

[0115] Further, it should be noted that the execution subject of the text recognition method can be a text recognition device. For example, the text recognition method can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the text recognition method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0116] Figure 5 is a block diagram of a text recognition device shown in an exemplary embodiment of the present application. As Figure 5 shown, the exemplary text recognition device 500 includes: a text prediction module 510, a mask generation module 520, a sampling module 530, a feature extraction module 540, and a text determination module 550. Specifically:

[0117] The text prediction module 510 is configured to perform text segmentation processing on the image to be recognized to obtain a segmentation image including a text region.

[0118] The mask generation module 520 is configured to perform binary marking processing on the segmentation image according to the pixel information of the text region to obtain an image feature mask.

[0119] The sampling module 530 is configured to perform sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampling feature.

[0120] The feature extraction module 540 is configured to input the sampling feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain a target image feature.

[0121] The text determination module 550 is configured to determine a corresponding target text according to the target image feature.

[0122] In this exemplary text recognition device, by performing text segmentation processing on the image to be recognized, a segmented image containing text regions can be obtained; by performing binary marking processing on the segmented image according to the pixel information of the text regions, an image feature mask with foreground and background separated can be obtained; by sampling the features to be recognized of the image to be recognized according to the image feature mask, the sampled features of the text regions in the image to be recognized can be sampled; by inputting the sampled features and the features to be recognized into a pre-trained attention network for feature extraction processing, the target image features output by the attention network can be obtained, and then the corresponding target text can be determined according to the target image features. Thus, the data processing amount of the neural network in the subsequent process can be reduced by the method of image sampling processing, and the accuracy of text recognition can also be improved.

[0123] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In actual application, the device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules to complete, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. No limitation is imposed here.

[0124] Among them, the functions of each module can be referred to in the text recognition method embodiment, and will not be repeated here.

[0125] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of an electronic device of the present application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is configured to execute program instructions stored in the memory 101 to implement the steps in any of the above text recognition method embodiments. In a specific implementation scenario, the electronic device 100 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 100 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited here.

[0126] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps in any of the above-described embodiments of the text recognition method. The processor 102 may also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 102 may be implemented jointly by integrated circuit chips.

[0127] In this exemplary electronic device, by performing text segmentation processing on the image to be recognized, a segmented image containing text regions can be obtained; by performing binary marking processing on the segmented image according to the pixel information of the text regions, an image feature mask with foreground and background separated can be obtained; by performing sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask, the sampling features of the text regions in the image to be recognized can be sampled; by inputting the sampling features and the feature to be recognized into a pre-trained attention network for feature extraction processing, the target image features output by the attention network can be obtained, and then the corresponding target text can be determined according to the target image features. Thus, the data processing amount of the neural network in the subsequent process can be reduced by the method of image sampling processing, and the accuracy of text recognition can also be improved.

[0128] Please refer to Figure 7 , Figure 7 is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 110 stores program instructions 111 that can be run by a processor, and the program instructions 111 are used to implement the steps in any of the above-described embodiments of the text recognition method.

[0129] In the exemplary storage medium, by running the program instructions in the storage medium, text segmentation processing is performed on the image to be recognized, and a segmented image including a text region can be obtained; binary marking processing is performed on the segmented image according to the pixel information of the text region to obtain an image feature mask with foreground and background separated; sampling processing is performed on the feature to be recognized of the image to be recognized according to the image feature mask, and then the sampling feature of the text region in the image to be recognized can be sampled; the sampling feature and the feature to be recognized are input into a pre-trained attention network for feature extraction processing, and the target image feature output by the attention network is obtained, and then the corresponding target text can be determined according to the target image feature. Thus, the data processing amount of the neural network in the subsequent process can be reduced by the method of image sampling processing, and the accuracy of text recognition can also be improved.

[0130] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be described in detail here.

[0131] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be described in detail in this article.

[0132] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0133] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

Claims

1. A text recognition method, characterized in that, The method includes: Performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region; Performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask; the image feature mask includes a foreground region and a background region; The step of performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask includes: performing sliding window statistical processing on the segmented image according to a preset window to obtain a statistical result; performing binary marking on the segmented image according to the statistical result to obtain an image feature mask; Performing sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampling feature; The step of performing sampling processing on the feature to be recognized of the image to be recognized according to the image feature mask to obtain a sampling feature includes: using the image feature mask as a sampling weight and determining the number of sampling times according to the number of foreground regions; performing discrete sampling processing on the feature to be recognized of the image to be recognized according to the number of sampling times to obtain the sampling feature; Inputting the sampling feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain a target image feature; The step of inputting the sampling feature and the feature to be recognized into a pre-trained attention network for feature extraction processing to obtain a target image feature includes: performing linear transformation processing on the sampling feature to obtain a query vector; performing linear transformation processing on the feature to be recognized to obtain a key vector and a value vector; performing feature extraction processing on the query vector, the key vector and the value vector through the attention network to obtain the target image feature; Determining a corresponding target text according to the target image feature.

2. The method according to claim 1, characterized in that Before performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region, the method further includes: Performing size adjustment processing on the received original image to obtain a to-be-processed image matching a preset size; Performing segmentation embedding processing on the to-be-processed image to obtain the image to be recognized and the feature to be recognized of the image to be recognized.

3. The method according to claim 1, characterized in that, The step of performing text segmentation processing on the image to be recognized to obtain a segmented image including a text region includes: Inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for segmentation processing and font standardization processing to obtain the segmented image; the font types of each text region in the segmented image are the same.

4. The method according to claim 3, characterized in that, The step of inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for segmentation processing and font standardization processing to obtain the segmented image includes: Inputting the feature to be recognized of the image to be recognized into a pre-trained image segmentation network for downsampling and feature extraction processing to obtain a downsampled feature of the feature to be recognized; Performing upsampling processing on the downsampled feature through the image segmentation network to obtain a segmented image including the text region.

5. The method according to claim 1, characterized in that The step of performing binary marking processing on the segmented image according to the pixel information of the text region to obtain an image feature mask includes: Partition the segmented image to obtain at least one image region; According to whether each pixel point in each image region is in the text region, respectively determine the binary label corresponding to each image region to obtain the image feature mask.

6. The method according to claim 5, wherein The segmented image includes the text region and the non-text region. Pixel points in the text region are foreground pixels, and pixel points in the non-text region are background pixels. The step of respectively determining the binary label corresponding to each image region according to whether each pixel point in each image region is in the text region to obtain the image feature mask includes: Obtain the foreground quantity proportion of the foreground pixels and the background quantity proportion of the background pixels in each image region; Perform first marking processing on the image regions where the foreground quantity proportion is greater than the background quantity proportion to obtain a first marking value; Perform second marking processing on the image regions where the foreground quantity proportion is less than or equal to the background quantity proportion to obtain a second marking value; the first marking value and the second marking value are different; Determine the image feature mask according to the first marking value and the second marking value.

7. The method according to claim 1, characterized in that The step of determining the corresponding target text according to the target image feature includes: According to the received prompt text, perform feature extraction processing on the prompt text to obtain prompt text features; Perform feature splicing processing on the prompt text features and the target image features to obtain spliced features; Input the spliced features into a pre-trained large language model to obtain the target text output by the large language model.

8. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method for recognizing natural scene text in any shape

    CN112183545A

  • Chinese handwritten text line identification method based on visual language joint reasoning

    CN115761764A