Text recognition method, electronic device and storage medium

By using an automated cropping and deduplication method for long image text recognition, the problem of low efficiency in long image text recognition is solved, and the integrity and consistency of text information are achieved.

CN115565175BActive Publication Date: 2025-10-28南京中孚信息技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211245543.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-12
Publication Date
2025-10-28
Estimated Expiration
2042-10-12

AI Technical Summary

Technical Problem

Existing technologies for long image text recognition are inefficient, requiring users to manually crop the text multiple times, which is time-consuming and laborious, and can easily lead to incomplete or repetitive text information recognition.

Method used

By automatically cropping the text image to be recognized according to the cropping size and order, the text detection boxes of the cropped image and the spliced ​​area image are determined, and deduplication is performed to finally sort and recognize the text information.

Benefits of technology

It enables automated cropping and recognition of text in long images, solving the problems of incomplete or repetitive text information and improving recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565175B_ABST
    Figure CN115565175B_ABST
Patent Text Reader

Abstract

This application provides a text recognition method, electronic device, and storage medium, relating to the field of data processing technology. First, it achieves automated cropping of the text image to be recognized. Then, for two adjacent cropped images, it determines the stitching region image, obtaining the corresponding document detection boxes for each stitched region image. This solves the problem that text at the connection points of adjacent cropped images may be truncated during image cropping, leading to incomplete or repetitive text information recognition. To address the issue of duplicate recognition, the text detection boxes corresponding to the cropped images can be further deduplicated using the text detection boxes corresponding to the stitched region images. Finally, by sorting the cropped images and stitched region images, and sequentially recognizing the document detection boxes corresponding to each cropped image and the text detection boxes corresponding to each stitched region image, the text information recognition result of the text image (long image) to be recognized is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a text recognition method, electronic device, and storage medium. Background Technology

[0002] Long images, as a type of image, are widely used because they can carry more information. A long image can be defined by the ratio of its longer side to its shorter side; generally, a ratio greater than or equal to 4 is considered a long image. Due to their larger size and the amount of text they contain, long images present greater challenges for text recognition.

[0003] In existing technologies, text recognition of long images requires users to manually crop the long image into multiple smaller images, and then recognize the text in each of the smaller images separately.

[0004] Because users need to manually crop the text multiple times to complete the recognition, it is time-consuming and laborious, resulting in low recognition efficiency for long image text. Summary of the Invention

[0005] The purpose of this application is to address the shortcomings of the prior art by providing a text recognition method, electronic device, and storage medium to solve the problem of low efficiency in long image recognition in the prior art.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, embodiments of this application provide a text recognition method, including:

[0008] Based on the cropping size of each cropping image corresponding to the text image to be identified and the cropping order of each cropping image, the text image to be identified is cropped sequentially to obtain a target number of cropped images arranged in order. At least one text detection box corresponding to each cropped image is determined sequentially. Each text detection box contains text information in the cropped image, and the text information contained in each text detection box does not overlap.

[0009] The stitching region image of each pair of adjacent cropped images is determined, as well as at least one text detection box corresponding to each stitched region image. Based on the at least one text detection box corresponding to each stitched region image, the text detection boxes corresponding to each cropped image are deduplicated.

[0010] The order of each cropped image and each spliced ​​region image is determined based on the cropping order of each cropped image in the text image to be identified, and the relationship between each spliced ​​region image and the cropped image.

[0011] According to the arrangement order of each cropped image and each stitched region image, the text information in each text detection box corresponding to each cropped image and the text information in each text detection box corresponding to each stitched region image are identified in turn, and the text information of each cropped image and the text information of each stitched region image are output in turn.

[0012] The text information is concatenated according to the output order of the text information of each cropped image and the text information of each spliced ​​region image to obtain the text information contained in the text image to be identified.

[0013] Optionally, the step of deduplicating the text detection boxes corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image includes:

[0014] The text detection boxes that overlap with the text detection boxes corresponding to each coordinate transformation of each cropped image and the text detection boxes corresponding to each coordinate transformation of each spliced ​​region image are identified as the first file detection boxes, thus obtaining the first detection box set.

[0015] The text detection boxes corresponding to each coordinate transformation of each stitched region image are determined as the second document detection boxes, thus obtaining the second detection box set;

[0016] Calculate the overlap index between each first document detection box in the first detection box set and each second document detection box in the second detection box set;

[0017] Based on the overlap index, duplicate text detection boxes corresponding to each cropped image after coordinate transformation are deduplicated.

[0018] Optionally, the stitched region image of each pair of adjacent cropped images and at least one text detection box corresponding to each stitched region image are determined, including:

[0019] Based on the text information detection results of two adjacent cropped images, the sub-cropping size of each cropped image in the two adjacent cropped images is determined, and sub-cropping images are cropped from the two adjacent cropped images according to the sub-cropping size of each cropped image in the two adjacent cropped images.

[0020] The sub-cropped images are stitched together to obtain the stitched area image corresponding to two adjacent cropped images;

[0021] A text detection algorithm is used to determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

[0022] Optionally, determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of the two adjacent cropped images includes:

[0023] If the text information detection result of the cropped image that is ranked first among two adjacent cropped images does not contain text information, then the sub-cropping size of the cropped image that is ranked first is determined according to the size of the cropped image that is ranked first.

[0024] If the text information detection result of the first cropped image in two adjacent cropped images contains text information, then the sub-cropping size of the first cropped image is determined based on the size of the first cropped image and the coordinates of the target text detection box corresponding to the first cropped image.

[0025] Optionally, determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of the two adjacent cropped images includes:

[0026] If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the cropped image that is ranked later is the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the size of the cropped image that is ranked last.

[0027] If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is not the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later.

[0028] Optionally, determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of the two adjacent cropped images includes:

[0029] If the text information detection result of the cropped image that is ranked later in two adjacent cropped images contains text information, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the coordinates of the target text detection box corresponding to the cropped image that is ranked later.

[0030] Optionally, before performing deduplication processing on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image, the method includes:

[0031] Based on the arrangement order of each cropped image and the arrangement order of each spliced ​​area image, determine the size deviation of each cropped image and the size deviation of each spliced ​​area image respectively.

[0032] Based on the size deviation of each cropped image, the coordinates of each text detection box corresponding to each cropped image are adjusted to obtain the text detection boxes corresponding to each cropped image after coordinate transformation.

[0033] Based on the size deviation of each stitched region image, the coordinates of each text detection box corresponding to each stitched region image are adjusted to obtain the text detection boxes corresponding to each stitched region image after coordinate transformation.

[0034] Optionally, before sequentially cropping the text image to be identified according to the cropping size and cropping order of each cropped image corresponding to the text image to be identified, to obtain a target number of cropped images arranged in sequence, the method includes:

[0035] Based on the size of the text image to be recognized, determine the ratio of the long side to the short side of the text image to be recognized;

[0036] Based on the ratio of the long side to the short side of the text image to be recognized, and the ratio of the long side to the short side of the largest text image that the text detection algorithm is allowed to recognize, the size of each image to be cropped and the number of images to be cropped are determined.

[0037] Secondly, embodiments of this application also provide a text recognition device, including: a determination module, a deduplication module, a text recognition module, and a text processing module;

[0038] The determining module is used to sequentially crop the text image to be identified according to the cropping size of each cropping image corresponding to the text image to be identified and the cropping order of each cropping image to be identified, to obtain a target number of cropped images arranged in sequence, and to sequentially determine at least one text detection box corresponding to each cropped image, wherein each text detection box contains text information in the cropped image, and the text information contained in each text detection box does not overlap.

[0039] The deduplication module is used to determine the splicing region image of each two adjacent cropped images and at least one text detection box corresponding to each splicing region image, and to perform deduplication processing on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each splicing region image.

[0040] The determining module is used to determine the arrangement order of each cropped image and each spliced ​​region image based on the cropping order of each cropped image in the text image to be identified, and the attribution relationship between each spliced ​​region image and the cropped image.

[0041] The text recognition module is used to sequentially recognize the text information in each text detection box corresponding to each cropped image and the text information in each text detection box corresponding to each spliced ​​region image according to the arrangement order of each cropped image and each spliced ​​region image, and sequentially output the text information of each cropped image and the text information of each spliced ​​region image.

[0042] The text processing module is used to concatenate the text information according to the output order of the text information of each cropped image and the text information of each spliced ​​region image to obtain the text information contained in the text image to be recognized.

[0043] Optionally, the deduplication module is specifically used to determine the text detection boxes that overlap with the text detection boxes corresponding to each cropped image after coordinate transformation and the text detection boxes corresponding to each spliced ​​region image as the first file detection boxes, thereby obtaining the first detection box set;

[0044] The text detection boxes corresponding to each coordinate transformation of each stitched region image are determined as the second document detection boxes, thus obtaining the second detection box set;

[0045] Calculate the overlap index between each first document detection box in the first detection box set and each second document detection box in the second detection box set;

[0046] Based on the overlap index, duplicate text detection boxes corresponding to each cropped image after coordinate transformation are deduplicated.

[0047] Optionally, the deduplication module is specifically used to determine the sub-cropping size of each cropped image in the two adjacent cropped images based on the text information detection results of the two adjacent cropped images, and to crop sub-cropping images from the two adjacent cropped images according to the sub-cropping size of each cropped image in the two adjacent cropped images.

[0048] The sub-cropped images are stitched together to obtain the stitched area image corresponding to two adjacent cropped images;

[0049] A text detection algorithm is used to determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

[0050] Optionally, the deduplication module is specifically used to determine the sub-cropping size of the first cropping image based on the size of the first cropping image if the text information detection result of the first cropping image in two adjacent cropping images does not contain text information.

[0051] If the text information detection result of the first cropped image in two adjacent cropped images contains text information, then the sub-cropping size of the first cropped image is determined based on the size of the first cropped image and the coordinates of the target text detection box corresponding to the first cropped image.

[0052] Optionally, the deduplication module is specifically used to determine the sub-cropping size of the cropping image that is sorted later in the order of two adjacent cropping images if the text information detection result of the cropping image that is sorted later does not contain text information and the sorting of the cropping image that is sorted later is the second to last.

[0053] If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is not the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later.

[0054] Optionally, the deduplication module is specifically used to determine the sub-cropping size of the cropping image that is ranked later in the order of two adjacent cropping images if the text information detection result of the cropping image ranked later contains text information, based on the size of the cropping image ranked later and the coordinates of the target text detection box corresponding to the cropping image ranked later.

[0055] Optionally, the device further includes: a coordinate transformation module;

[0056] The coordinate transformation module is specifically used to determine the size deviation of each cropped image and the size deviation of each spliced ​​area image according to the arrangement order of each cropped image and the arrangement order of each spliced ​​area image.

[0057] Based on the size deviation of each cropped image, the coordinates of each text detection box corresponding to each cropped image are adjusted to obtain the text detection boxes corresponding to each cropped image after coordinate transformation.

[0058] Based on the size deviation of each stitched region image, the coordinates of each text detection box corresponding to each stitched region image are adjusted to obtain the text detection boxes corresponding to each stitched region image after coordinate transformation.

[0059] Optionally, the determining module is further configured to determine the ratio of the long side to the short side of the text image to be recognized based on the size of the text image to be recognized;

[0060] Based on the ratio of the long side to the short side of the text image to be recognized, and the ratio of the long side to the short side of the largest text image that the text detection algorithm is allowed to recognize, the size of each image to be cropped and the number of images to be cropped are determined.

[0061] Thirdly, embodiments of this application provide an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method provided in the first aspect.

[0062] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the method provided in the first aspect.

[0063] The beneficial effects of this application are:

[0064] The text recognition method, electronic device, and storage medium provided in this application firstly crop the text image to be recognized sequentially according to the cropping size and cropping order of the images to be cropped, obtaining a target number of cropped images arranged in sequence, thus achieving automated cropping of the text image to be recognized. For each cropped image, the corresponding document detection box is identified. Then, for two adjacent cropped images, the stitching region image is determined, and the corresponding document detection box is obtained, thus solving the problem that the text at the connection point of adjacent cropped images may be truncated during image cropping, resulting in incomplete or repetitive text information recognition. To address the problem of repeated recognition, the text detection boxes corresponding to the cropped images can be further deduplicated using the text detection boxes corresponding to the stitched region images. Finally, by sorting the cropped images and stitched region images, and sequentially recognizing the document detection boxes corresponding to the cropped images and the text detection boxes corresponding to the stitched region images, the text information recognition result of the text image (long image) to be recognized is obtained. Attached Figure Description

[0065] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 Flowchart of the text recognition method provided in the embodiments of this application Figure 1 ;

[0067] Figure 2 Flowchart of the text recognition method provided in the embodiments of this application Figure 2 ;

[0068] Figure 3Flowchart of the text recognition method provided in the embodiments of this application Figure 3 ;

[0069] Figure 4 Flowchart of the text recognition method provided in the embodiments of this application Figure 4 ;

[0070] Figure 5 A schematic diagram of a coordinate system provided for an embodiment of this application;

[0071] Figure 6 Flowchart of the text recognition method provided in the embodiments of this application Figure 5 ;

[0072] Figure 7 Flowchart of the text recognition method provided in the embodiments of this application Figure 6 ;

[0073] Figure 8 Flowchart of the text recognition method provided in the embodiments of this application Figure 7 ;

[0074] Figure 9 A schematic diagram of a text recognition device provided in an embodiment of this application;

[0075] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0077] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0078] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0079] First, a brief explanation of the relevant background information for this application will be provided.

[0080] Long images are widely used because their larger size allows them to convey more information; for example, in the new media industry, information is disseminated through long images. This has led to a demand for OCR (Optical Character Recognition) processing of these long images.

[0081] However, existing OCR recognition systems often do not support OCR for long images. On the one hand, long images are large in size and contain a lot of text, resulting in long processing times and difficulty in quickly responding to user requests, as well as placing a heavy burden on the hardware. On the other hand, in order to reduce computational pressure, OCR recognition often scales the image. When the long side of the image is scaled to a fixed size, because the ratio of the long side to the short side of the long image is too large, scaling the short side proportionally will result in the text in the image being too small and losing a lot of image information, leading to poor recognition results.

[0082] Therefore, existing OCR recognition systems either do not support OCR recognition of long images, or require users to manually select the recognition area of ​​the image, crop the image to a suitable size, and then perform OCR recognition. If OCR recognition of the entire long image is required, the long image needs to be cropped into multiple smaller images, and then OCR recognition of each smaller image needs to be performed separately.

[0083] The method of manually cropping long images before performing OCR recognition requires users to crop multiple times to meet the recognition requirements, which is time-consuming and labor-intensive and cannot achieve automation. In addition, if the cropping size is fixed, the text may be truncated at the connection point of the image crop, resulting in problems such as misrecognition, missed recognition, or duplicate recognition.

[0084] Based on this, the text recognition method provided in this application firstly crops the long image as a whole according to a pre-set cropping size to obtain multiple cropped images, thereby achieving automated cropping of the long image. For cases where text is truncated at the connection points of adjacent cropped images, sub-cropping is performed on two adjacent cropped images to obtain an image of the spliced ​​region between the two adjacent cropped images. Since there may be overlap between the spliced ​​region image and the cropped image, the text detection boxes corresponding to the cropped images can be deduplicated using the text detection boxes corresponding to the spliced ​​region images. Finally, by sorting the cropped images and spliced ​​region images, and sequentially recognizing the file detection boxes corresponding to each cropped image and the text detection boxes corresponding to each spliced ​​region image, the text information recognition result of the long image is obtained.

[0085] Figure 1 Flowchart of the text recognition method provided in the embodiments of this application Figure 1 The subject executing this method can be a computer device. For example... Figure 1 As shown, the method may include:

[0086] S101. Based on the cropping size of each cropping image corresponding to the text image to be identified and the cropping order of each cropping image, the text image to be identified is cropped sequentially to obtain the target number of cropped images arranged in order, and at least one text detection box corresponding to each cropped image is determined sequentially. Each text detection box contains text information in the cropped image, and the text information contained in each text detection box does not overlap.

[0087] The text image to be recognized refers to a long image, meaning the image size meets preset conditions. There isn't a formal definition of a long image; here, it's defined by the ratio of the longer side to the shorter side. Generally, a ratio greater than or equal to 4 is considered a long image. For simplicity, the following explanation focuses on the case where the longer side is the image's height and the shorter side is its width. The method is similar for images where the longer side is the width and the shorter side is the height.

[0088] It's important to clarify that the image to be cropped is not the actual image. The text image to be recognized is cropped according to the cropping dimensions of the image to be cropped, resulting in a cropped image. For example, if the text image to be recognized is 5 cm high and 2 cm wide, and the cropping dimensions are set to crop the text image from the 2 cm high position along a cropping line parallel to the width, then we will obtain a cropped image and a remaining image. The cropped image will be 2 cm high and 2 cm wide, and the remaining image will be 3 cm high and 2 cm wide. The remaining image can be further cropped according to the cropping dimensions of the image to be cropped until the entire text image to be recognized is cropped, resulting in multiple cropped images.

[0089] In addition, there is a cropping order when cropping images. The cropping order of the images to be cropped determines the order of the cropped images. In this embodiment, when cropping the text image to be recognized, it is cropped in order from top to bottom, similar to cropping a piece of white paper from top to bottom.

[0090] After cropping the text image to be recognized according to the cropping size of the image to be cropped, the first cropped image is obtained. The remaining image in the text image to be recognized is then cropped again according to the cropping size of the first cropped image, resulting in the second cropped image, and so on, until the cropped images are arranged in order. Since the cropping size of the image to be cropped is fixed, the number of cropped images obtained is also fixed.

[0091] Of course, in actual cropping, the cropping size of each image can be different.

[0092] For each cropped image obtained, a text detection box can be determined for the text information contained in each cropped image. The usual text detection algorithm is to use the text detection box to circle the text information and to recognize the text in the text detection box.

[0093] S102. Determine the stitching region image of each pair of adjacent cropped images, and at least one text detection box corresponding to each stitched region image, and perform deduplication processing on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image.

[0094] Given a fixed cropping size, for two adjacent cropped images, since the text information at the connection point between the two adjacent cropped images may be truncated during image cropping, this embodiment can further process each pair of adjacent cropped images to obtain a spliced ​​region image of the two adjacent cropped images. The spliced ​​region image will contain the text information that may have been truncated at the connection point between the two adjacent cropped images.

[0095] Similarly, at least one text detection box can be determined for each stitched region image. Since the text information contained in the stitched region image may be duplicated with the text information contained in its two adjacent cropped images, the text detection boxes corresponding to each cropped image can be deduplicated based on the at least one text detection box corresponding to each stitched region image.

[0096] S103. Determine the arrangement order of each cropped image and each spliced ​​region image based on the cropping order of each cropped image in the text image to be recognized and the attribution relationship between each spliced ​​region image and the cropped image.

[0097] The cropping order of each cropped image represents the arrangement order of each cropped image in the text image to be recognized, and the arrangement order of each cropped image can be determined based on the cropping order of each cropped image.

[0098] The relationship between the stitched region image and the cropped image can refer to which two adjacent cropped images the stitched region image is obtained by processing. Assuming that the stitched region image is the stitched region image corresponding to cropped image 1 and cropped image 2, then the text information contained in the stitched region image is actually located between the text information contained in cropped image 1 and the text information contained in cropped image 2. Therefore, the order of the stitched region images can also be determined.

[0099] Therefore, the cropped images and the stitched region images can be sorted as a whole to obtain the sorting result.

[0100] S104. According to the arrangement order of each cropped image and each stitched region image, the text information in each text detection box corresponding to each cropped image and the text information in each text detection box corresponding to each stitched region image are identified in sequence, and the text information of each cropped image and the text information of each stitched region image are output in sequence.

[0101] Optionally, text information can be recognized for each image sequentially according to the arrangement order of each cropped image and each stitched region image. For a specific image (cropped image or stitched region image), the sorting order of the corresponding text detection boxes can be: for the text detection boxes corresponding to each row in the image, they can be arranged in order from left to right, and for each row, they can be arranged in order from top to bottom.

[0102] Based on the sorting order of each cropped image and each stitched image, as well as the arrangement order of each text detection box in each cropped image and each stitched image, the text information in each text detection box corresponding to each cropped image and the text information in each text detection box corresponding to each stitched region image can be identified in sequence according to the arrangement order, and the identified text information can be output in sequence.

[0103] Here, once the text information in a file detection box is recognized, it can be directly output. That is, after recognizing each file detection box, a corresponding text information is output, and the output text information is the text information contained in the file detection box that has just been recognized.

[0104] S105. The text information is concatenated according to the output order of the text information of each cropped image and the text information of each spliced ​​region image to obtain the text information contained in the text image to be recognized.

[0105] In some application scenarios, when it is desired to obtain the complete text information of the text image to be recognized, the text information can be concatenated in the order of the output of the recognized text information, that is, chained together, to obtain the text information contained in the text image to be recognized.

[0106] In summary, the text recognition method provided in this embodiment first sequentially crops the text image to be recognized according to the cropping size and cropping order of the images to be cropped, obtaining a target number of cropped images arranged in sequence, thus achieving automated cropping of the text image to be recognized. For each cropped image, the corresponding document detection box is identified. Then, for two adjacent cropped images, the stitching region image is determined, and the corresponding document detection box is obtained, thus solving the problem that the text at the connection point of adjacent cropped images may be truncated during image cropping, resulting in incomplete or repetitive text information recognition. To address the problem of repeated recognition, the text detection boxes corresponding to the cropped images can be further deduplicated using the text detection boxes corresponding to the stitched region images. Finally, by sorting the cropped images and stitched region images, and sequentially identifying the document detection boxes corresponding to the cropped images and the text detection boxes corresponding to the stitched region images, the text information recognition result of the text image (long image) to be recognized is obtained.

[0107] Figure 2 Flowchart of the text recognition method provided in the embodiments of this application Figure 2 Optionally, in step S102, the process of deduplicating the text detection boxes corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image may include:

[0108] S201. The text detection boxes that overlap with the text detection boxes corresponding to the coordinate transformation of each cropped image and the text detection boxes corresponding to the coordinate transformation of each spliced ​​region image are identified as the first file detection boxes, thus obtaining the first detection box set.

[0109] In one feasible approach, based on the text detection boxes corresponding to each cropped image and each text detection box corresponding to each stitched region image obtained above, the coordinates of the text detection boxes can be determined. This allows for the identification of text detection boxes in each cropped image whose coordinates fall within the stitched region image, and the combination of these text detection boxes yields a first set of detection boxes, denoted as A = {a}. 1 ,a 2 ,…,a i ,…}.

[0110] It should be noted that the coordinates of the file detection boxes corresponding to each cropped image obtained above are the coordinates of the file detection boxes in the cropped image. However, when performing deduplication on the text detection boxes, it is necessary to convert the coordinates of the file detection boxes corresponding to each cropped image into coordinates in the text image to be recognized, and to convert the coordinates of the file detection boxes corresponding to each spliced ​​region image into coordinates in the text image to be recognized.

[0111] In one implementation, when no text information is detected in any of the stitched region images, that is, the stitched region images do not contain text information, then all text detection boxes in the first detection box set can be retained.

[0112] Otherwise, you can continue with steps S202-S204.

[0113] S202. Determine the text detection boxes corresponding to each coordinate transformation of each stitched region image as the second document detection boxes, and obtain the second detection box set.

[0114] Optionally, a second set of detection boxes can be formed from the text detection boxes corresponding to the coordinate transformations of the stitched region images obtained above, which can be denoted as set B = {b}. 1 ,b 2 ,...,b i ,…}.

[0115] S203. Calculate the overlap index of each first document detection box in the first detection box set and each second document detection box in the second detection box set.

[0116] Optionally, the IOU (intersection over union) can be calculated for each text detection box in set B and each text detection box in set A. IOU is the ratio of the area of ​​the intersection to the area of ​​the union of two text detection boxes, and it is a method to determine whether text boxes overlap. Assume the coordinates of any text detection box in set B are... The coordinates of any text detection box in set A are: So

[0117]

[0118] For each element in set B, an IOU is calculated with each element in set A. For example, b 1 and a 1 Calculation yields b 1 and a 1 The corresponding IOU value; b 2 and a 2 Calculation yields b 2 and a 2The corresponding IOU value. The IOU value calculated for any two text detection boxes is used as their overlap index.

[0119] S204. Based on the overlap index, deduplicate the text detection boxes corresponding to each cropped image after coordinate transformation.

[0120] If the IOU of the text detection boxes in set B with those in set A is 0, then these text detection boxes with IOU = 0 are retained. If there are text detection boxes in set B with IOU greater than 0 with those in set A, further judgment is required. If IOU > 0.3, then the two text detection boxes are considered to belong to the same text. Here, 0.3 is an empirical value that can be adjusted according to the actual situation. In this case, the one with the larger area is retained, and the other is deleted. If the areas of the two text detection boxes are the same, then the corresponding text detection box in set B is retained, and the corresponding text detection box in set A is deleted.

[0121] Figure 3 Flowchart of the text recognition method provided in the embodiments of this application Figure 3 Optionally, in step S102, determining the stitched region image of each pair of adjacent cropped images and at least one text detection box corresponding to each stitched region image may include:

[0122] S301. Based on the text information detection results of two adjacent cropped images, determine the sub-cropping size of each cropped image in the two adjacent cropped images respectively, and crop the sub-cropping images from the two adjacent cropped images according to the sub-cropping size of each cropped image in the two adjacent cropped images respectively.

[0123] This embodiment describes a method for generating a stitched region image from adjacent cropped images.

[0124] Extract the first and second cropped images. The first and second cropped images are then a pair of adjacent cropped images. Then take the second and third cropped images. The second and third cropped images are then a pair of adjacent cropped images. Repeat this process until the (n-1)th and nth cropped images are taken.

[0125] Optionally, the sub-cropping size of the two adjacent cropped images can be determined based on the text information detection results of the two adjacent cropped images, so that a sub-cropping image can be cropped from each of the two adjacent cropped images according to the sub-cropping size.

[0126] S302. Stitch the sub-cropped images together to obtain the stitched area image corresponding to two adjacent cropped images.

[0127] Here, the images can be stitched together according to the order in which adjacent cropped images are arranged, with the sub-cropped image cropped from the later cropped image being stitched after the sub-cropped image cropped from the earlier cropped image.

[0128] S303. Using a text detection algorithm, determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

[0129] For each stitched region image obtained, text detection can also be performed to determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

[0130] Figure 4 Flowchart of the text recognition method provided in the embodiments of this application Figure 4 Optionally, in step S101, before sequentially cropping the text image to be recognized according to the cropping size of each cropping image corresponding to the text image to be recognized and the cropping order of each cropping image to be recognized, to obtain the target number of cropped images arranged in sequence, the method of this application may further include:

[0131] S401. Determine the ratio of the long side to the short side of the text image to be recognized based on the size of the text image to be recognized.

[0132] Assuming the dimensions of the text image to be recognized are denoted as W (width, short side) and H (height, long side), then the aspect ratio is denoted as r, where r = H ÷ W.

[0133] S402. Based on the ratio of the long side to the short side of the text image to be recognized, and the ratio of the long side to the short side of the largest text image that the text detection algorithm is allowed to recognize, determine the size of each image to be cropped and the number of images to be cropped.

[0134] Let R be the maximum aspect ratio that the text detection algorithm can recognize. R is a constant because the algorithm's capabilities are limited. It can only recognize images with an aspect ratio of R at most. In practical use, factors such as computing power and processing time can be considered to reduce the value of R.

[0135] In one case, when r ≤ R, it means that for the current ORC recognition system, the text image to be recognized is not a long image and does not need to be cropped.

[0136] Figure 5 This is a schematic diagram of a coordinate system provided in an embodiment of this application. In this embodiment, when describing the coordinates of an image, the coordinates are based on the image's position within a coordinate system. Figure 5 The coordinates are in the coordinate system shown.

[0137] In another case, when r > R, the image cropping step of this method can be executed. The number of cropped images obtained by cropping is denoted as n, n = ceil(r ÷ R), where ceil refers to rounding up. The width w' of each cropped image obtained after cropping is w' = W. Except for the last cropped image (the cropped image sorted last), the height of the cropped images is h' = ceil(H ÷ n), and the height h n of the last cropped image is h = H - (n - 1) × ceil(H ÷ n). The height of the i-th cropped image is also denoted as h i (i = 1, 2, 3…, n).

[0138] Based on this, in step S101, according to the cropping sizes of each cropped image corresponding to the text image to be recognized and the cropping order of each cropped image, sequentially cropping the text image to be recognized to obtain the target number of cropped images arranged in sequence may include:

[0139] If the cropping order of the current cropped image is not the last one (i < n), then according to the cropping order of the current cropped image and the cropping size of the current cropped image, determine the cropping area of the current cropped image.

[0140] If the cropping order of the current cropped image is the last one (i = n), then according to the cropping size of the current cropped image and the number of cropped images, determine the cropping area of the current cropped image.

[0141] That is, when cropping an image, a four-tuple coordinate c = (x1, y1, x2, y2) (which also refers to the cropping size of the cropped image) can be used to represent the cropping position (cropping area). (x1, y1) represents the coordinates of the upper left corner point of the image, and (x2, y2) represents the coordinates of the lower right corner point of the image. When i < n, the cropping of the i-th cropped image is performed according to the cropping area c = (0, (i - 1) × h', w', i × h'); when i = n, the cropping of the n-th cropped image is performed according to the cropping area c = (0, (n - 1) × h', w', H).

[0142] Optionally, based on the determined cropping area, the cropping area can be cropped from the text image to be recognized to obtain a cropped image.

[0143] When using a text detection algorithm to detect the position of text in each cropped image, it can be to first detect at least one file detection box corresponding to the cropped image. The position of the j-th text detection box in the i-th cropped image can be represented by an eight-tuple coordinate Let (x1, y1) represent the coordinates of the top-left corner of the detection box, (x2, y2) represent the coordinates of the top-right corner of the detection box, (x3, y3) represent the coordinates of the bottom-right corner of the detection box, and (x4, y4) represent the coordinates of the bottom-left corner of the detection box. The coordinates of the first text detection box are... The coordinates of the last text detection box are

[0144] Figure 6 Flowchart of the text recognition method provided in the embodiments of this application Figure 5 In step S301, based on the text information detection results of two adjacent cropped images, the sub-cropping size of each cropped image in the two adjacent cropped images is determined, which may include:

[0145] S601. If the text information detection result of the cropped image that is ranked first among two adjacent cropped images does not contain text information, then the sub-cropping size of the cropped image that is ranked first is determined according to the size of the cropped image that is ranked first.

[0146] In this embodiment, when generating the stitched region image, the sub-cropping size in two adjacent cropped images is determined.

[0147] Assuming that the two adjacent cropped images extracted are the i-th cropped image and the (i+1)-th cropped image, then the cropped image that is sorted first refers to the i-th cropped image.

[0148] For the i-th cropped image, if no text information is detected, then the sub-cropping size of the cropped image is c = (0, 0.75 × h′, w′, h′), where 0.75 is an empirical number that can be adjusted according to the actual situation.

[0149] S602. If the text information detection result of the cropped image that is ranked first among two adjacent cropped images contains text information, then the sub-cropping size of the cropped image that is ranked first is determined according to the size of the cropped image that is ranked first and the coordinates of the target text detection box corresponding to the cropped image that is ranked first.

[0150] In this embodiment, the target text detection box corresponding to the first cropped image can refer to the last text detection box corresponding to the first cropped image.

[0151] If text information is detected, the coordinates of the last text detection box at this point are... If ((y3+y4)÷2)>0.9×h′, where 0.9 is an empirical number that can be adjusted according to the actual situation, then the sub-cutting size c=(0,min(y1,y2),w′,h′), where min means taking the minimum value between the two; otherwise, the sub-cutting size c=(0,max(max(y3,y4),0.75×h′),w′,h′), where max means taking the maximum value between the two, where 0.75 is an empirical number that can be adjusted according to the actual situation.

[0152] Figure 7 Flowchart of the text recognition method provided in the embodiments of this application Figure 6 In step S301, based on the text information detection results of two adjacent cropped images, the sub-cropping size of each cropped image in the two adjacent cropped images is determined, which may include:

[0153] S701. If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the size of the cropped image that is ranked last.

[0154] Based on the above example, the cropped image that is sorted later refers to the (i+1)th cropped image.

[0155] For the (i+1)th cropped image, if no text information is detected, and when i = n-1 (i.e., the cropped image that follows is the second to last in the order), then the sub-crop size is c = (0, 0, w′, min(0.25 × h′, h n )), min means taking the minimum value between the two.

[0156] S702. If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is not the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later.

[0157] If no text information is detected in the (i+1)th cropped image, and if i is not equal to n-1 (the cropped images that follow are not the second to last in the order), then the sub-crop size is c = (0, 0, w′, 0.25 × h′), where 0.25 is an empirical number that can be adjusted according to the actual situation.

[0158] S703. If the text information detection result of the cropped image that is ranked later in two adjacent cropped images contains text information, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the coordinates of the target text detection box corresponding to the cropped image that is ranked later.

[0159] In this embodiment, the target text detection box corresponding to the cropped image that is sorted later refers to the first text detection box corresponding to the cropped image that is sorted later.

[0160] When text information is detected in the (i+1)th cropped image, the coordinates of the first text detection box are... If ((y1+y2)÷2)<0.1×h i+1 h i Let be the height of the (i+1)th cropped image, where 0.1 is an empirical number that can be adjusted according to the actual situation. Then, the sub-cropping size c = (0, 0, max(y3, y4), w′, h i+1 Otherwise, the sub-cutting size c = (0, min(min(y1, y2), 0.25 × h′), w′, h i+1 The value of 0.25 is an empirical figure and can be adjusted according to the actual situation.

[0161] Figure 8 Flowchart of the text recognition method provided in the embodiments of this application Figure 7 In step S102, before performing deduplication processing on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image, the method of this application may further include:

[0162] S801. Based on the arrangement order of each cropped image and the arrangement order of each spliced ​​area image, determine the size deviation of each cropped image and the size deviation of each spliced ​​area image respectively.

[0163] Assuming cropped image 1 is arranged first and cropped image 2 is arranged second, the size deviation of the first cropped image can generally be considered 0. Starting with the second cropped image, the size deviation of each cropped image can be the sum of the heights of all the cropped images preceding it. Determining the size deviation of the stitched region image is similar; however, for the stitched region image, the size deviation needs to be calculated by combining the heights of the cropped images preceding it and the height of the stitched region image itself.

[0164] S802. Based on the size deviation of each cropped image, adjust the coordinates of each text detection box corresponding to each cropped image to obtain the text detection box after coordinate transformation for each cropped image.

[0165] Optionally, based on the size deviation of each cropped image calculated in the above steps, the text detection box after coordinate transformation of each cropped image can be calculated. This can be achieved by adding the size deviation of the cropped image to the ordinates of the four points of the text detection box in the cropped image.

[0166] S803. Based on the size deviation of each spliced ​​region image, adjust the coordinates of each text detection box corresponding to each spliced ​​region image to obtain the text detection boxes after coordinate transformation corresponding to each spliced ​​region image.

[0167] Similarly, based on the size deviation of each stitched region image calculated in the above steps, the text detection box after coordinate transformation of each stitched region image can be calculated.

[0168] Optionally, the text recognition algorithm can be invoked sequentially to recognize the text information in each file detection box of each cropped image, according to the arrangement order of each cropped image and each stitched region image.

[0169] It is worth noting that the method of this application can be flexibly embedded into various OCR algorithms. Generally speaking, OCR algorithms can be divided into single-stage algorithms and two-stage algorithms. A single-stage algorithm can obtain the position of the text detection box and the text information contained in the text detection box at the same time each OCR algorithm is called. A two-stage algorithm can first detect the text detection box in the image through a text detection algorithm, and then recognize the text information in the text detection box through a text recognition algorithm.

[0170] The difference between the two methods lies in the order of steps and efficiency. For the two-stage algorithm, the steps are: first, calculate the parameter r; then, crop the text image to be recognized to obtain multiple cropped images; apply a text detection algorithm to all cropped images; then, obtain the stitched region image of two adjacent cropped images; and again call the text detection algorithm to detect whether there is text information in the stitched region image. After removing duplicate text detection boxes from the cropped images and stitched region images, finally, perform text recognition according to the arrangement order of the cropped images and stitched region images, and the arrangement order of the text detection boxes in the cropped images and stitched region images, to obtain the text recognition result.

[0171] For a single-stage algorithm, the steps are as follows: first, calculate the parameter r; then, crop the text image to be recognized to obtain multiple cropped images; for all cropped images, use the OCR algorithm to simultaneously obtain text detection boxes and text information; then, based on the obtained stitched region image of two adjacent cropped images, call the OCR algorithm again to obtain text detection boxes and text information from the stitched region image; after removing duplicate text detection boxes, the text recognition result can be obtained.

[0172] Generally, the detection time of text detection boxes is much shorter than the text recognition time. Therefore, in this application, the text detection boxes of all cropped images and all spliced ​​region images are obtained first, and then text recognition is performed on the sorted text detection boxes to avoid the detection of redundant text detection boxes, which would increase the text recognition time.

[0173] In summary, the text recognition method provided in this embodiment first sequentially crops the text image to be recognized according to the cropping size and cropping order of the images to be cropped, obtaining a target number of cropped images arranged in sequence, thus achieving automated cropping of the text image to be recognized. For each cropped image, the corresponding document detection box is identified. Then, for two adjacent cropped images, the stitching region image is determined, and the corresponding document detection box is obtained, thus solving the problem that the text at the connection point of adjacent cropped images may be truncated during image cropping, resulting in incomplete or repetitive text information recognition. To address the problem of repeated recognition, the text detection boxes corresponding to the cropped images can be further deduplicated using the text detection boxes corresponding to the stitched region images. Finally, by sorting the cropped images and stitched region images, and sequentially identifying the document detection boxes corresponding to the cropped images and the text detection boxes corresponding to the stitched region images, the text information recognition result of the text image (long image) to be recognized is obtained.

[0174] The following describes the apparatus, device, and storage medium used to execute the text recognition method provided in this application. The specific implementation process and technical effects are described above and will not be repeated below.

[0175] Figure 9 This is a schematic diagram of a text recognition device provided in an embodiment of this application. The function implemented by this text recognition device corresponds to the steps performed by the method described above. This device can be understood as the aforementioned computer equipment, such as... Figure 9 As shown, the device may include: a determination module 910, a deduplication module 920, a text recognition module 930, and a text processing module 940;

[0176] The determining module 910 is used to sequentially crop the text image to be identified according to the cropping size of each cropping image corresponding to the text image to be identified and the cropping order of each cropping image to be identified, to obtain a target number of cropped images arranged in sequence, and to sequentially determine at least one text detection box corresponding to each cropped image, wherein each text detection box contains text information in the cropped image, and the text information contained in each text detection box does not overlap.

[0177] The deduplication module 920 is used to determine the splicing region image of each of two adjacent cropped images and at least one text detection box corresponding to each splicing region image, and to perform deduplication processing on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each splicing region image.

[0178] The determining module 910 is used to determine the arrangement order of each cropped image and each spliced ​​region image based on the cropping order of each cropped image in the text image to be recognized and the attribution relationship between each spliced ​​region image and the cropped image.

[0179] The text recognition module 930 is used to sequentially recognize the text information in the text detection boxes corresponding to each cropped image and each spliced ​​region image according to the arrangement order of each cropped image and each spliced ​​region image, and sequentially output the text information of each cropped image and each spliced ​​region image.

[0180] The text processing module 940 is used to concatenate the text information according to the output order of the text information of each cropped image and the text information of each spliced ​​region image to obtain the text information contained in the text image to be recognized.

[0181] Optionally, the deduplication module 920 is specifically used to determine the text detection boxes that overlap with the text detection boxes corresponding to each cropped image after coordinate transformation and the text detection boxes corresponding to each spliced ​​region image as the first file detection boxes, so as to obtain the first detection box set.

[0182] The text detection boxes corresponding to each coordinate transformation of each stitched region image are determined as the second document detection boxes, thus obtaining the second detection box set;

[0183] Calculate the overlap index between each first document detection box in the first detection box set and each second document detection box in the second detection box set;

[0184] Based on the overlap index, duplicate text detection boxes corresponding to each cropped image after coordinate transformation are deduplicated.

[0185] Optionally, the deduplication module 920 is specifically used to determine the sub-cropping size of each cropped image in the two adjacent cropped images based on the text information detection results of the two adjacent cropped images, and to crop sub-cropping images from the two adjacent cropped images according to the sub-cropping size of each cropped image in the two adjacent cropped images.

[0186] The sub-cropped images are stitched together to obtain the stitched area image corresponding to two adjacent cropped images;

[0187] A text detection algorithm is used to determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

[0188] Optionally, the deduplication module 920 is specifically used to determine the sub-cropping size of the first cropping image based on the size of the first cropping image if the text information detection result of the first cropping image in two adjacent cropping images does not contain text information.

[0189] If the text information detection result of the first cropped image in two adjacent cropped images contains text information, then the sub-cropping size of the first cropped image is determined based on the size of the first cropped image and the coordinates of the target text detection box corresponding to the first cropped image.

[0190] Optionally, the deduplication module 920 is specifically used to determine the sub-cropping size of the cropping image that is sorted later in the two adjacent cropping images if the text information detection result of the cropping image that is sorted later does not contain text information and the sorting of the cropping image that is sorted later is the second to last.

[0191] If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is not the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later.

[0192] Optionally, the deduplication module 920 is specifically used to determine the sub-cropping size of the cropping image that is ranked later in the order of two adjacent cropping images if the text information detection result of the cropping image ranked later contains text information, based on the size of the cropping image ranked later and the coordinates of the target text detection box corresponding to the cropping image ranked later.

[0193] Optionally, the device further includes: a coordinate transformation module;

[0194] The coordinate transformation module is specifically used to determine the size deviation of each cropped image and the size deviation of each stitched area image based on the arrangement order of each cropped image and the arrangement order of each stitched area image.

[0195] Based on the size deviation of each cropped image, the coordinates of each text detection box corresponding to each cropped image are adjusted to obtain the text detection boxes corresponding to each cropped image after coordinate transformation.

[0196] Based on the size deviation of each stitched region image, the coordinates of each text detection box corresponding to each stitched region image are adjusted to obtain the text detection boxes corresponding to each stitched region image after coordinate transformation.

[0197] Optionally, the determining module 910 is further configured to determine the ratio of the long side to the short side of the text image to be recognized based on the size of the text image to be recognized;

[0198] Based on the ratio of the long side to the short side of the text image to be recognized, and the ratio of the long side to the short side of the largest text image that the text detection algorithm is allowed to recognize, the size of each image to be cropped and the number of images to be cropped are determined.

[0199] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0200] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0201] The modules described above can be connected or communicate with each other via wired or wireless connections. Wired connections can include metal cables, optical fibers, hybrid cables, or any combination thereof. Wireless connections can include connections via LAN, WAN, Bluetooth, ZigBee, or NFC, or any combination thereof. Two or more modules can be combined into a single module, and any module can be divided into two or more units. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here.

[0202] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The device may be a computing device with data processing capabilities.

[0203] The device may include: processor 801 and memory 802.

[0204] The memory 802 is used to store programs, and the processor 801 calls the programs stored in the memory 802 to execute the above method embodiments. The specific implementation and technical effects are similar, and will not be described again here.

[0205] The memory 802 stores program code, which, when executed by the processor 801, causes the processor 801 to perform various steps in the methods according to various exemplary embodiments of this application described in the "Exemplary Methods" section above.

[0206] The processor 801 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0207] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0208] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs the above-described method embodiments.

[0209] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0210] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0211] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.

[0212] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A text recognition method, characterized in that, include: Based on the cropping size of each cropping image corresponding to the text image to be identified and the cropping order of each cropping image, the text image to be identified is cropped sequentially to obtain a target number of cropped images arranged in order. At least one text detection box corresponding to each cropped image is determined sequentially. Each text detection box contains text information in the cropped image, and the text information contained in each text detection box does not overlap. The stitching region image of each pair of adjacent cropped images is determined, as well as at least one text detection box corresponding to each stitched region image. Based on the at least one text detection box corresponding to each stitched region image, the text detection boxes corresponding to each cropped image are deduplicated. The order of each cropped image and each spliced ​​region image is determined based on the cropping order of each cropped image in the text image to be identified, and the relationship between each spliced ​​region image and the cropped image. According to the arrangement order of each cropped image and each stitched region image, the text information in each text detection box corresponding to each cropped image and the text information in each text detection box corresponding to each stitched region image are identified in turn, and the text information of each cropped image and the text information of each stitched region image are output in turn. The text information is concatenated according to the output order of the text information of each cropped image and the text information of each spliced ​​region image to obtain the text information contained in the text image to be recognized. The step of deduplicating the text detection boxes corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image includes: The text detection boxes that overlap with the text detection boxes corresponding to each coordinate transformation of each cropped image and the text detection boxes corresponding to each coordinate transformation of each spliced ​​region image are identified as the first file detection boxes, thus obtaining the first detection box set. The text detection boxes corresponding to each coordinate transformation of each stitched region image are determined as the second document detection boxes, thus obtaining the second detection box set; Calculate the overlap index of each first document detection box in the first detection box set and each second document detection box in the second detection box set; the overlap index is used to characterize the ratio of the area of ​​the intersection of the first document detection box and the second document detection box to the area of ​​the union of the first document detection box and the second document detection box. Based on the overlap index, duplicate text detection boxes corresponding to each cropped image after coordinate transformation are deduplicated.

2. The method according to claim 1, characterized in that, Determine the stitched region image of each pair of adjacent cropped images, and at least one text detection box corresponding to each stitched region image, including: Based on the text information detection results of two adjacent cropped images, the sub-cropping size of each cropped image in the two adjacent cropped images is determined, and sub-cropping images are cropped from the two adjacent cropped images according to the sub-cropping size of each cropped image in the two adjacent cropped images. The sub-cropped images are stitched together to obtain the stitched area image corresponding to two adjacent cropped images; A text detection algorithm is used to determine at least one text detection box corresponding to the stitched region image of two adjacent cropped images.

3. The method according to claim 2, characterized in that, The step of determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of two adjacent cropped images includes: If the text information detection result of the cropped image that is ranked first among two adjacent cropped images does not contain text information, then the sub-cropping size of the cropped image that is ranked first is determined according to the size of the cropped image that is ranked first. If the text information detection result of the first cropped image in two adjacent cropped images contains text information, then the sub-cropping size of the first cropped image is determined based on the size of the first cropped image and the coordinates of the target text detection box corresponding to the first cropped image.

4. The method according to claim 2, characterized in that, The step of determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of two adjacent cropped images includes: If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the cropped image that is ranked later is the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the size of the cropped image that is ranked last. If the text information detection result of the cropped image that is ranked later in two adjacent cropped images does not contain text information, and the ranking of the cropped image that is ranked later is not the second to last, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later.

5. The method according to claim 2, characterized in that, The step of determining the sub-cropping size of each cropped image in two adjacent cropped images based on the text information detection results of two adjacent cropped images includes: If the text information detection result of the cropped image that is ranked later in two adjacent cropped images contains text information, then the sub-cropping size of the cropped image that is ranked later is determined according to the size of the cropped image that is ranked later and the coordinates of the target text detection box corresponding to the cropped image that is ranked later.

6. The method according to claim 1, characterized in that, Before performing deduplication on each text detection box corresponding to each cropped image based on at least one text detection box corresponding to each stitched region image, the method includes: Based on the arrangement order of each cropped image and the arrangement order of each spliced ​​area image, determine the size deviation of each cropped image and the size deviation of each spliced ​​area image respectively. Based on the size deviation of each cropped image, the coordinates of each text detection box corresponding to each cropped image are adjusted to obtain the text detection boxes corresponding to each cropped image after coordinate transformation. Based on the size deviation of each stitched region image, the coordinates of each text detection box corresponding to each stitched region image are adjusted to obtain the text detection boxes corresponding to each stitched region image after coordinate transformation.

7. The method according to claim 1, characterized in that, Before sequentially cropping the text image to be identified according to the cropping size and cropping order of each cropping image corresponding to the text image to be identified, to obtain a target number of cropped images arranged in sequence, the method includes: Based on the size of the text image to be recognized, determine the ratio of the long side to the short side of the text image to be recognized; Based on the ratio of the long side to the short side of the text image to be recognized, and the ratio of the long side to the short side of the largest text image that the text detection algorithm is allowed to recognize, the size of each image to be cropped and the number of images to be cropped are determined.

8. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the text recognition method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, performs the steps of the text recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • PDF document table extraction method, device and equipment and computer readable storage medium

    CN110390269A

  • Image-text processing method and device, and display method and device, equipment and storage medium

    CN112882678A