Text recognition method, device, electronic device and storage medium

Through the dual watermarking algorithm and lightweight recognition network, the problem of poor watermark image recognition effect is solved, efficient and accurate text recognition is achieved, and labor costs are reduced.

CN114581646BActive Publication Date: 2025-07-11SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111485442.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-07-11
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

When processing images with watermarks, the existing text recognition algorithm has poor recognition effect, high recognition error rate, low recognition accuracy and efficiency, and high labor cost.

Method used

The double watermarking algorithm is used to detect the watermark type and remove the watermark through the watermark detection network, and the text detection network is used to extract text boxes, combine the convolutional neural network and recurrent neural network for feature extraction and transcription, and build a lightweight recognition network to improve recognition accuracy and efficiency.

Benefits of technology

It significantly improves the recognition accuracy and efficiency of watermark images, reduces labor costs, and improves the recognition accuracy and efficiency of ordinary images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581646B_ABST
    Figure CN114581646B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides a text recognition method, apparatus, electronic device, and storage medium. The method includes: using a watermark detection network to detect a to-be-recognized image, obtaining a watermark type and a watermark detection box, and selecting a watermark removal model, using the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image; using a text detection network to detect text in the watermark-free image, obtaining the position of a text box in the watermark-free image, cropping the text box based on the position of the text box to obtain a text box; using the text box as the input of a text recognition network, using a convolutional neural network layer to extract features from the text box to obtain a first feature map, and using a recurrent neural network layer to process the first feature map to obtain a second feature map, using a transcription layer to transcribe the second feature map to obtain the text in the to-be-recognized image. This disclosure improves the accuracy and precision of text recognition, and at the same time has a high text recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a text recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of computer technologies, methods for using text recognition technologies to recognize text in images have been widely used. However, in some application scenarios, there may be some watermarks in the images in addition to the text. The watermarks in the images are likely to obscure the text. Additionally, the text in some images is relatively dense and the text content in the images is relatively complex. These situations will all have an adverse impact on the results of text recognition in the images.

[0003] In traditional text recognition algorithms, taking the text recognition in engineering certificates as an example, in order to avoid the above situations from affecting the recognition results, manual extraction and verification are required, thus increasing the time cost of the staff. And when the certificate text is relatively dense and the content is relatively complex, it is very easy to have recognition errors. When applying existing OCR recognition algorithms for text recognition, the problems of watermark occlusion and text blurring cannot be solved either. Therefore, when existing text recognition algorithms are used to recognize text in watermark images, there are problems such as poor recognition effect, high recognition error rate, and low recognition accuracy and efficiency.

[0004] In view of the above problems in the prior art, it is necessary to provide a text recognition method for watermark images with high recognition accuracy and efficiency and low labor cost. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a text recognition method, apparatus, electronic device, and storage medium to solve the problems of high recognition error rate, poor recognition accuracy, and low efficiency in text recognition of watermark images existing in the prior art.

[0006] In a first aspect of embodiments of the present disclosure, a text recognition method is provided, including: obtaining an image to be recognized, detecting the image to be recognized by using a watermark detection network to obtain a watermark type and a watermark detection frame, and determining a watermark removal model matching the watermark type, using the watermark removal model to remove the watermark in the watermark detection frame to obtain a watermark-free image; performing a text detection operation on the watermark-free image by using a text detection network to obtain the positions of text frames in the watermark-free image, cropping the text frames based on the positions of the text frames to obtain text frames; using the text frames as the input of a text recognition network, extracting features from the text frames by using a convolutional neural network layer to obtain a first feature map, and processing the first feature map by using a recurrent neural network layer to obtain a second feature map, and transcribing the second feature map by using a transcription layer to obtain the text in the image to be recognized.

[0007] In a second aspect of the embodiments of the present disclosure, there is provided a table structure extraction device, including: a watermark detection module configured to obtain an image to be recognized, detect the image to be recognized by using a watermark detection network to obtain a watermark type and a watermark detection box, determine a watermark removal model matching the watermark type, and use the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image; a text detection module configured to perform text detection on the watermark-free image by using a text detection network to obtain the positions of text boxes in the watermark-free image, and crop the text boxes based on the positions of the text boxes to obtain text boxes; a text recognition module configured to use the text boxes as inputs to a text recognition network, extract features from the text boxes by using a convolutional neural network layer to obtain a first feature map, process the first feature map by using a recurrent neural network layer to obtain a second feature map, and transcribe the second feature map by using a transcription layer to obtain the text in the image to be recognized.

[0008] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.

[0009] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above method.

[0010] The above at least one technical solution adopted in the embodiments of the present disclosure can achieve the following beneficial effects:

[0011] By obtaining the image to be recognized, detecting the image to be recognized by using a watermark detection network to obtain a watermark type and a watermark detection box, determining a watermark removal model matching the watermark type, and using the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image; performing text detection on the watermark-free image by using a text detection network to obtain the positions of text boxes in the watermark-free image, and cropping the text boxes based on the positions of the text boxes to obtain text boxes; using the text boxes as inputs to a text recognition network, extracting features from the text boxes by using a convolutional neural network layer to obtain a first feature map, processing the first feature map by using a recurrent neural network layer to obtain a second feature map, and transcribing the second feature map by using a transcription layer to obtain the text in the image to be recognized. The present disclosure not only has high recognition accuracy and efficiency for ordinary watermark-free images, but also has high recognition accuracy and recognition precision for watermarked images, and improves the efficiency of text recognition of watermarked images and reduces the time cost of personnel. Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 is a schematic flowchart of the text recognition method provided by the embodiments of the present disclosure;

[0014] Figure 2 is a schematic flowchart of the processing of the dual watermark removal algorithm provided by the embodiments of the present disclosure;

[0015] Figure 3 is a schematic structural diagram of the text recognition device provided by the embodiments of the present disclosure;

[0016] Figure 4 is a schematic structural diagram of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0017] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0018] As mentioned above, in real-world scenarios, in addition to the text to be recognized in the image to be recognized, there will also be some watermarks, such as seals and transparent watermarks. The watermarks in the image are likely to obscure the text. When recognizing the text obscured by the watermark, it will greatly increase the difficulty of text recognition. In addition, the text in some images to be recognized is relatively dense, and the text content in the image to be recognized is also relatively complex. These situations have an adverse impact on the recognition results of the text in the image to be recognized. Taking the text recognition in the fields of commercial certificates and engineering certificates as an example, the problems of the existing text recognition algorithms and the improvement points of the present disclosure will be described in detail below, which can specifically include the following content:

[0019] When identifying the product name and merchant information in engineering certificates, since there are often a large number of seals and watermarks in engineering certificate documents, this will affect the results of text recognition. Traditional text recognition algorithms mainly include the following two aspects: First, manual extraction and verification are required, which not only increases the time cost of staff, but also easily leads to recognition errors when the certificate text is dense and the content is complex; Second, existing OCR recognition algorithms are used for text recognition, but existing OCR recognition algorithms cannot well solve the problems of watermark occlusion and text blurring. Therefore, directly applying existing OCR recognition algorithms often fails to achieve ideal results.

[0020] In view of the above considerations of existing technical problems, the present disclosure proposes a dual watermark removal algorithm to accurately and efficiently remove the seals and transparent watermarks existing in the picture. Through this processing, the recognition accuracy of the model can be significantly improved. In addition, in order to ensure the recognition speed in dense scenarios, a lightweight recognition network is proposed, which has a smaller model size and higher recognition efficiency.

[0021] Figure 1 It is a schematic flow chart of the text recognition method provided by an embodiment of the present disclosure. Figure 1 The text recognition method can be executed by a server. As Figure 1 shown, the text recognition method may specifically include:

[0022] S101, obtain the image to be recognized, use the watermark detection network to detect the image to be recognized, obtain the watermark type and the watermark detection box, and determine the watermark removal model matching the watermark type, and use the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image;

[0023] S102, use the text detection network to perform text detection operations on the watermark-free image, obtain the positions of the text boxes in the watermark-free image, and crop the text boxes based on the positions of the text boxes to obtain the text boxes;

[0024] S103, use the text box as the input of the text recognition network, use the convolutional neural network layer to extract features from the text box to obtain the first feature map, and use the recurrent neural network layer to process the first feature map to obtain the second feature map, and use the transcription layer to transcribe the second feature map to obtain the text in the image to be recognized.

[0025] Specifically, the image to be recognized can be an object in the form of a picture generated according to a business certificate, an engineering certificate, etc. The image to be recognized can be either an image containing a seal and / or a watermark, or an image without a watermark. For an image without a watermark, traditional OCR recognition algorithms can be directly used for text recognition. For an image containing a seal and / or a watermark, it is necessary to first use a watermark detection network to determine the type of watermark in the image to be recognized, and then select an appropriate watermark removal algorithm based on the watermark type to remove the watermark.

[0026] Furthermore, the present disclosure proposes two watermark removal algorithms, which can be understood as a dual watermark removal algorithm, namely a seal removal algorithm and a transparent watermark removal algorithm. The watermark detection network can not only identify the specific position of the watermark, but also detect the type of watermark. If the detection result only contains a seal, it is processed by the seal removal algorithm. If the detection result only contains a transparent watermark, it is processed by the transparent watermark removal algorithm. If the detection result shows that both a seal and a transparent watermark exist in the image to be recognized, it is processed successively by the seal removal algorithm and the transparent watermark removal algorithm. These processes can significantly improve the text recognition accuracy.

[0027] According to the technical solution provided by the embodiments of the present disclosure, the present disclosure obtains the image to be recognized, uses a watermark detection network to detect the image to be recognized, obtains the watermark type and the watermark detection frame, determines a watermark removal model matching the watermark type, uses the watermark removal model to remove the watermark in the watermark detection frame to obtain a watermark-free image; uses a text detection network to perform a text detection operation on the watermark-free image to obtain the position of the text box in the watermark-free image, and crops the text box based on the position of the text box to obtain the text box; takes the text box as the input of the text recognition network, uses a convolutional neural network layer to extract features from the text box to obtain a first feature map, and uses a recurrent neural network layer to process the first feature map to obtain a second feature map, and uses a transcription layer to transcribe the second feature map to obtain the text in the image to be recognized. The present disclosure not only has high recognition accuracy and efficiency for ordinary watermark-free images, but also has high recognition accuracy and recognition precision for watermarked images, and improves the efficiency of text recognition of watermarked images and reduces the personnel time cost.

[0028] In some embodiments, using a watermark detection network to detect the image to be recognized, obtaining the watermark type and the watermark detection frame includes: taking the image to be recognized as the input of the watermark detection network, using the watermark detection network to detect the image to be recognized, so as to judge the watermark type in the image to be recognized based on the convolutional layer in the watermark detection network, and generate a watermark detection frame for the watermark; wherein, the watermark detection network is a neural network model composed of a convolutional layer using a normalization network and an activation function, and the watermark type includes a seal and a transparent watermark.

[0029] Specifically, the present disclosure constructs a watermark detection network model. Using the watermark detection network model, not only can the type of watermark in the image to be recognized (i.e., belonging to a seal or a transparent watermark) be detected, but also the specific position of the watermark (i.e., the watermark detection box) can be obtained. The watermark detection network model is a convolutional neural network. For example, it can be a neural network model composed of convolutional layers with normalization networks and activation functions. The input of the watermark detection network model is the image to be recognized. After a series of convolutional operations, the output result is the watermark type and the watermark detection box, that is, whether the image to be recognized contains a watermark, which category the watermark belongs to, and the judgment result of the watermark position. The output of the watermark detection network model is used as the input of the watermark removal model, and the watermark removal model is used to remove the watermark.

[0030] In some embodiments, determining a watermark removal model that matches the watermark type includes: when the watermark in the image to be recognized is a seal, using a preset seal removal model to remove the seal in the image to be recognized; when the watermark in the image to be recognized is a transparent watermark, using a preset transparent watermark removal model to remove the seal in the image to be recognized; when the watermark in the image to be recognized is both a seal and a transparent watermark, sequentially using the preset seal removal model and transparent watermark removal model to remove the seal and the transparent watermark in the image to be recognized.

[0031] Specifically, the watermark removal model includes two types, namely a seal removal model and a transparent watermark removal model. Among them, the seal removal model is used to remove the seal in the image to be recognized, and the transparent watermark removal model is used to remove the transparent watermark in the image to be recognized. These two types of watermark removal models proposed by the present disclosure together constitute a dual watermark removal algorithm. Below, in conjunction with the accompanying drawings, the processing flow of the dual watermark removal algorithm will be described in detail. Figure 2 It is a schematic diagram of the processing flow of the dual watermark removal algorithm provided by an embodiment of the present disclosure. Figure 2 The processing flow of the dual watermark removal algorithm specifically may include the following contents:

[0032] Input the image to be recognized, use the watermark detection network to detect the watermark, obtain the watermark type and the watermark detection box in the image to be recognized, and then three processing branches are generated. The content of the first processing branch is that when only a transparent watermark is included in the watermark detection result, the transparent watermark removal algorithm is used to remove the transparent watermark to obtain a watermark-free image; the content of the second processing branch is that when only a seal is included in the watermark detection result, the seal removal algorithm is used to remove the seal to obtain a watermark-free image; the content of the third processing branch is that when both a transparent watermark and a seal are included in the watermark detection result, the seal and the watermark are sequentially removed through the seal removal algorithm and the transparent watermark removal algorithm to obtain a watermark-free image.

[0033] It should be noted that the reason for choosing multi-branch processing in this disclosure is that the processing time of the stamp removal algorithm is short, while the transparent watermark removal algorithm requires encoding and decoding stages, which will bring relatively high time consumption. Through data analysis, it is found that in the actual scenario, the images to be recognized often contain a large number of stamp occlusions and only a small number of transparent watermark occlusions. Therefore, by designing the processing logic into a multi-branch rather than a serial structure, not only the recognition accuracy of the algorithm is improved, but also the recognition speed is enhanced.

[0034] In some embodiments, a preset stamp removal model is used to remove the stamps in the image to be recognized, including: converting the image to be recognized from the RGB color space to the HSV color space, and judging the color of the stamp according to the chroma value of each pixel in the watermark detection frame, obtaining the color layer corresponding to the color of the stamp, and expanding the color layer into three channels to obtain a watermark-free image.

[0035] Specifically, the stamp removal model adopts a stamp removal algorithm based on color space conversion, that is, by converting the original picture (i.e., the image to be recognized) from the RGB color space to the HSV color space, and then judging whether the H value of each pixel in the stamp area is within the threshold range, that is, judging the color of the current stamp through the chroma value, such as a blue stamp, a red stamp or a green stamp. Then, the layer corresponding to the current stamp color is obtained and expanded into three channels. At this time, the picture after removing the stamp can be obtained.

[0036] In some embodiments, a preset transparent watermark removal model is used to remove the watermarks in the image to be recognized, including: sequentially using the encoder and decoder in the transparent watermark removal model to perform encoding and decoding operations on the image to be recognized, where the encoding operation is used to perform encoding calculation on the image feature information in the image to be recognized to remove the pixel points corresponding to the watermarks in the image to be recognized, and obtain a watermark-free image.

[0037] Specifically, the transparent watermark removal model consists of two parts: an encoder and a decoder. The transparent watermark removal model is a neural network model trained based on a watermark picture training set obtained through preprocessing. The working principle of the transparent watermark removal model is to use the encoder to obtain the picture feature information of each watermark, and then use the decoder to process the original picture into a watermark-free image based on the picture feature information.

[0038] Further, when using the decoder to remove the watermarks in the watermark detection frame, the decoder decodes and calculates the image feature information in the feature map, judges the similarity between the pixel points in each watermark detection frame and other background information according to the image feature information, judges the probability value of which pixel point is a watermark pixel point based on the similarity, and removes the data of the pixel points according to the probability value.

[0039] Further, before training the transparent watermark removal model, a batch of watermarked images and non-watermarked images can be generated artificially. In practical applications, the source dataset can be obtained by first acquiring non-watermarked images and then adding watermarks to the non-watermarked images. Considering that watermarks in real-world scenarios may exist in any position in the image in various forms, such distribution changes are also simulated when generating training data. For example, the position and size of the watermark are randomly generated in the non-watermarked image. Eventually, the model continuously learns this mapping during the training process to enable it to have the ability to remove watermarks.

[0040] In some embodiments, performing a text detection operation on the non-watermarked image using a text detection network to obtain the positions of text boxes in the non-watermarked image includes: processing the non-watermarked image using a feature extraction network in the text detection network to obtain a feature map, predicting the probability values corresponding to the feature map to obtain a probability map, superimposing the probability map with a threshold map to obtain a new feature map, where the positions of the text boxes are included in the new feature map, and taking the positions of the text boxes and the confidence scores as the output of the text detection network.

[0041] Specifically, after removing the watermark in the image to be recognized, performing text detection on the non-watermarked image using a text detection network to obtain the text boxes in the non-watermarked image; the feature extraction network in the text detection network performs convolution and upsampling operations on the non-watermarked image, and then uses two branches to process the convolutional feature map respectively to obtain a probability map and a threshold map, superimposing the probability map and the threshold map together to obtain a new feature map, and based on the positions of the text boxes in the new feature map, extracting several text boxes from the feature map.

[0042] Further, after cropping out the text boxes based on their positions, taking the cropped text boxes as the input of the text recognition network. That is to say, the input of the text recognition network is the text boxes obtained by performing text detection on the non-watermarked image. In practical applications, the text detection network can adopt the DBnet network.

[0043] In some embodiments, a convolutional neural network layer is used to extract features from a text box to obtain a first feature map, and a recurrent neural network layer is used to process the first feature map to obtain a second feature map. The transcription layer transcribes the second feature map to obtain the text in the image to be recognized, including: performing a convolution operation on the text box using the backbone network in the convolutional neural network layer, and inputting the feature map output by the backbone network into consecutive depthwise hybrid convolution blocks. Channels with different convolutional kernels in the depthwise hybrid convolution blocks are used to perform convolution on the feature map output by the backbone network, and the first feature map obtained by convolution is used as the input of the recurrent neural network layer; the second feature map obtained by the recurrent neural network layer processing the first feature map is used as the input of the transcription layer, so that the transcription layer performs a transcription operation on the second feature map to obtain the text in the image to be recognized.

[0044] Specifically, in the process of certificate recognition in the case of dense text, it is necessary to ensure the recognition speed. Therefore, the present disclosure constructs a lightweight recognition network as the Backbone (feature extraction) in the recognition stage. The specific structure and processing process of the lightweight recognition network will be described in detail below in combination with specific embodiments, which may specifically include the following contents:

[0045] The lightweight recognition network structure mainly consists of three parts, namely, a convolutional neural network layer, a recurrent neural network layer, and a transcription layer. Among them, the convolutional neural network layer includes a Stem backbone network. In the backbone network, 3 3x3 convolutional kernels are used to replace the traditional 7x7 convolutional kernel. In this way, while ensuring the receptive field, the network parameters and computational amount are reduced. After being processed by the Stem stage, the feature map is input into the middle layer of the network. The middle layer of the network includes four consecutive stages, and each stage is stacked by 2 depthwise hybrid convolution blocks MixConv block. Depthwise hybrid convolution MixConv is a convolutional method improved on layer-by-layer depthwise separable convolution and grouped convolution. Depthwise hybrid convolution MixConv groups the input feature map by channels, and then applies convolutional kernels of different sizes on each group. For example, the following three convolutional kernels can be used: 3x3 convolutional kernel, 5x5 convolutional kernel, 7x7 convolutional kernel, where the 5x5 convolutional kernel and 7x7 convolutional kernel are realized by convolving 3x3 convolutional kernels with different dilation rates. In this way, the network parameters can be reduced, and depthwise hybrid convolution MixConv can capture information of different scales inside the convolutional kernel, which is beneficial to extracting richer features.

[0046] Further, after the first feature map is obtained by processing with the convolutional neural network layer, the first feature map is input into the recurrent neural network layer for processing to obtain a second feature map, and finally the transcription layer transcribes the second feature map output by the recurrent neural network into text, so as to obtain the text in the final image to be recognized.

[0047] It should be noted that in the present disclosure, the ReLU activation function in the network is replaced by H-Swish. Compared with ReLU, the Swish activation function can significantly improve the network accuracy. However, because its calculation process includes exponential operations, the calculation efficiency is not high. Therefore, the present disclosure uses H-Swish to replace Swish, and its function formula is as follows;

[0048]

[0049] In addition, the present disclosure introduces the SE attention module into the network, and at the same time replaces the Sigmoid activation function in the original SE module with H-Sigmoid. Similar to Swish, the Sigmoid activation function also contains exponential operations. Therefore, the present disclosure uses its accelerated version H-Sigmoid to replace it. The SE attention module weights the channels, emphasizes the effective information, and suppresses the invalid information, significantly improving the network accuracy while incurring a very small computational cost.

[0050] According to the technical solution provided by the embodiments of the present disclosure, the present disclosure proposes an algorithm for recognizing text in commercial certificates and engineering certificates, which can automatically extract key fields in the certificates using an AI model, avoiding misrecognition caused by humans in dense scenarios while reducing costs and improving efficiency. In addition, for the problem that watermarks have a great impact on the recognition accuracy, before text recognition, a watermark detection module is used to obtain the watermark in the text, and then the proposed dual watermark removal algorithm is used to remove the seal and transparent watermark in the picture, thus avoiding the influence of watermarks on recognition. Finally, based on the research results of MobileNet V3, a lightweight recognition network is constructed, which uses depthwise separable convolutions to extract multi-scale features inside the convolutional kernels, and at the same time changes the activation function in the network, making the extracted features more abundant, thereby improving the recognition accuracy. Based on the above-mentioned processing, not only the error rate of text recognition is reduced, but also the text recognition accuracy and efficiency are improved. The technical solution of the present disclosure has a good recognition result for watermarked images.

[0051] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For the details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.

[0052] Figure 3 is a schematic structural diagram of the text recognition device provided by the embodiment of the present disclosure. As Figure 3 shown, the text recognition device includes:

[0053] The watermark detection module 301 is configured to obtain an image to be recognized, detect the image to be recognized using a watermark detection network, obtain the watermark type and the watermark detection box, determine a watermark removal model matching the watermark type, and use the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image;

[0054] The text detection module 302 is configured to perform a text detection operation on the watermark-free image using a text detection network, obtain the positions of the text boxes in the watermark-free image, and crop the text boxes based on the positions of the text boxes to obtain the text boxes;

[0055] The text recognition module 303 is configured to use the text box as the input of the text recognition network, extract features from the text box using a convolutional neural network layer to obtain a first feature map, process the first feature map using a recurrent neural network layer to obtain a second feature map, and transcribe the second feature map using a transcription layer to obtain the text in the image to be recognized.

[0056] In some embodiments, Figure 3 The watermark detection module 301 takes the image to be recognized as the input of the watermark detection network, and uses the watermark detection network to detect the image to be recognized, so as to judge the watermark type in the image to be recognized based on the convolutional layer in the watermark detection network, and generate a watermark detection box for the watermark; wherein, the watermark detection network is a neural network model composed of convolutional layers using a normalization network and an activation function, and the watermark types include seals and transparent watermarks.

[0057] In some embodiments, Figure 3 When the watermark in the image to be recognized is a seal, the watermark removal module 304 uses a preset seal removal model to remove the seal in the image to be recognized; when the watermark in the image to be recognized is a transparent watermark, the watermark removal module 304 uses a preset transparent watermark removal model to remove the seal in the image to be recognized; when the watermark in the image to be recognized is a seal and a transparent watermark, the preset seal removal model and the transparent watermark removal model are used in sequence to remove the seal and the transparent watermark in the image to be recognized.

[0058] In some embodiments, Figure 3 The watermark removal module 304 converts the image to be recognized from the RGB color space to the HSV color space, judges the color of the seal according to the chroma value of each pixel in the watermark detection box, obtains a color layer corresponding to the color of the seal, and expands the color layer into three channels to obtain a watermark-free image.

[0059] In some embodiments, Figure 3The watermark removal module 304 sequentially utilizes the encoder and decoder in the transparent watermark removal model to perform an encoding operation and a decoding operation on the image to be recognized. The encoding operation is used to perform encoding calculations on the image feature information in the image to be recognized, so as to remove the pixel points corresponding to the watermark in the image to be recognized and obtain a watermark-free image.

[0060] In some embodiments, Figure 3 The text detection module 302 of the watermark-free image processes the watermark-free image by using the feature extraction network in the text detection network to obtain a feature map, predicts the probability values corresponding to the feature map to obtain a probability map, superimposes the probability map and the threshold map to obtain a new feature map. The position of the text box is included in the new feature map, and the position of the text box and the confidence score are used as the output of the text detection network.

[0061] In some embodiments, Figure 3 The text recognition module 303 of the watermark-free image performs a convolution operation on the text box by using the backbone network in the convolutional neural network layer, and inputs the feature map output by the backbone network into a continuous deep hybrid convolutional block. Channels with different convolutional kernels in the deep hybrid convolutional block are used to perform convolution on the feature map output by the backbone network, and the first feature map obtained by convolution is used as the input of the recurrent neural network layer; the second feature map obtained by the recurrent neural network layer processing the first feature map is used as the input of the transcription layer, so that the transcription layer performs a transcription operation on the second feature map to obtain the text in the image to be recognized.

[0062] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0063] Figure 4 is a schematic structural diagram of the electronic device 4 provided by the embodiments of the present disclosure. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.

[0064] Exemplarily, the computer program 403 may be divided into one or more modules / units. One or more modules / units are stored in the memory 402 and executed by the processor 401 to complete the present disclosure. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 403 in the electronic device 4.

[0065] The electronic device 4 can be a desktop computer, a notebook, a handheld computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 4 These are merely examples of the electronic device 4 and do not constitute a limitation on the electronic device 4. It may include more or fewer components than shown in the figure, or combine some components, or have different components. For example, the electronic device may also include input / output devices, network access devices, a bus, etc.

[0066] The processor 401 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0067] The memory 402 can be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 can also include both the internal storage unit and the external storage device of the electronic device 3. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 can also be used to temporarily store data that has been output or will be output.

[0068] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0069] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0070] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0071] In the embodiments provided by this disclosure, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.

[0072] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0073] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0074] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above method embodiments of the present disclosure may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0075] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A text recognition method, characterized in that, Including: Obtain the image to be recognized, detect the image to be recognized using a watermark detection network to obtain the watermark type and the watermark detection box, and determine a watermark removal model that matches the watermark type, and use the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image; Perform text detection on the watermark-free image using a text detection network to obtain the position of the text box in the watermark-free image, and crop the text box based on the position of the text box to obtain the text box; Use the text box as the input of the text recognition network, use the convolutional neural network layer to extract features from the text box to obtain a first feature map, and use the recurrent neural network layer to process the first feature map to obtain a second feature map, and use the transcription layer to transcribe the second feature map to obtain the text in the image to be recognized; Among them, the determining the watermark removal model that matches the watermark type includes: When the watermark in the image to be recognized is a seal, use a preset seal removal model to remove the seal in the image to be recognized; When the watermark in the image to be recognized is a transparent watermark, use a preset transparent watermark removal model to remove the transparent watermark in the image to be recognized; When the watermark in the image to be recognized is a seal and a transparent watermark, sequentially use a preset seal removal model and a transparent watermark removal model to remove the seal and the transparent watermark in the image to be recognized; The using the convolutional neural network layer to extract features from the text box to obtain a first feature map, and using the recurrent neural network layer to process the first feature map to obtain a second feature map, and using the transcription layer to transcribe the second feature map to obtain the text in the image to be recognized includes: Use the backbone network in the convolutional neural network layer to perform a convolution operation on the text box, and input the feature map output by the backbone network into consecutive depth hybrid convolution blocks, use the channels with different convolution kernels in the depth hybrid convolution blocks to perform convolution on the feature map output by the backbone network, and use the convolved first feature map as the input of the recurrent neural network layer; Use the second feature map obtained by the recurrent neural network layer processing the first feature map as the input of the transcription layer, so that the transcription layer performs a transcription operation on the second feature map to obtain the text in the image to be recognized.

2. The method according to claim 1, characterized in that, The using the watermark detection network to detect the image to be recognized to obtain the watermark type and the watermark detection box includes: Use the image to be recognized as the input of the watermark detection network, and use the watermark detection network to detect the image to be recognized, so as to judge the watermark type in the image to be recognized based on the convolutional layer in the watermark detection network and generate the watermark detection box of the watermark; Among them, the watermark detection network is a neural network model composed of convolutional layers using a normalization network and an activation function, and the watermark type includes a seal and a transparent watermark.

3. The method according to claim 1, wherein The using the preset seal removal model to remove the seal in the image to be recognized includes: Convert the image to be recognized from the RGB color space to the HSV color space, and determine the color of the seal according to the chroma value of each pixel in the watermark detection box. Obtain the color layer corresponding to the color of the seal, and expand the color layer into three channels to obtain a watermark-free image.

4. The method according to claim 1, wherein The removal of the transparent watermark in the image to be recognized by using a preset transparent watermark removal model includes: Successively use the encoder and decoder in the transparent watermark removal model to perform encoding and decoding operations on the image to be recognized. The encoding operation is used to perform encoding calculations on the image feature information in the image to be recognized to remove the pixel points corresponding to the watermark in the image to be recognized, and obtain a watermark-free image.

5. The method according to claim 1, wherein The execution of text detection operations on the watermark-free image by using a text detection network to obtain the positions of the text boxes in the watermark-free image includes: Use the feature extraction network in the text detection network to process the watermark-free image to obtain a feature map, predict the probability values corresponding to the feature map to obtain a probability map, superimpose the probability map and the threshold map to obtain a new feature map. The position of the text box is included in the new feature map. The position of the text box and the confidence score are used as the output of the text detection network.

6. A text recognition device, characterized in that, Including: A watermark detection module, configured to obtain an image to be recognized, use a watermark detection network to detect the image to be recognized, obtain the watermark type and the watermark detection box, determine a watermark removal model matching the watermark type, and use the watermark removal model to remove the watermark in the watermark detection box to obtain a watermark-free image; A text detection module, configured to use a text detection network to perform text detection operations on the watermark-free image, obtain the positions of the text boxes in the watermark-free image, and crop the text boxes based on the positions of the text boxes to obtain the text boxes; A text recognition module, configured to use the text box as the input of a text recognition network, use a convolutional neural network layer to extract features from the text box to obtain a first feature map, use a recurrent neural network layer to process the first feature map to obtain a second feature map, and use a transcription layer to transcribe the second feature map to obtain the text in the image to be recognized; Among them, a watermark removal module is further included, which is used to, when the watermark in the image to be recognized is a seal, use a preset seal removal model to remove the seal in the image to be recognized; when the watermark in the image to be recognized is a transparent watermark, use a preset transparent watermark removal model to remove the transparent watermark in the image to be recognized; When the watermark in the image to be recognized is a seal and a transparent watermark, successively use a preset seal removal model and a transparent watermark removal model to remove the seal and the transparent watermark in the image to be recognized; The text recognition module is used to perform a convolution operation on the text box by using the backbone network in the convolutional neural network layer, and input the feature map output by the backbone network into consecutive deep hybrid convolution blocks. Channels with different convolution kernels in the deep hybrid convolution blocks are used to perform convolution on the feature map output by the backbone network, and the first feature map obtained by convolution is used as the input of the recurrent neural network layer; the second feature map obtained by the recurrent neural network layer processing the first feature map is used as the input of the transcription layer, so that the transcription layer performs a transcription operation on the second feature map to obtain the text in the image to be recognized.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Court judgment document-oriented multi-scale learning character recognition method and system

    CN111985464A

  • Document identification method and identification system

    CN113205049A