Image text detection method, device, storage medium and electronic device
By extracting text areas in image text detection, text recognition and determining the writing order, the problem that the recognition results in the prior art do not conform to ordinary people's reading comprehension is solved, and an image text output that is easier to understand is achieved.
Patent Information
- Application Number
- CN202111250649.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-10-26
AI Technical Summary
The prior art cannot effectively determine the writing order of text in image text detection, resulting in the recognition results that do not conform to the order of ordinary people's reading comprehension, affecting users to quickly understand the meaning of the image.
By extracting the text area from the image to be detected, text recognition is performed, text writing order is determined, and the identification text is adjusted according to the order to output sequential copy that conforms to ordinary people's reading comprehension.
It realizes the adjustment of the recognition results according to the detected text writing order, breaking through the mechanical output problem of traditional detection, and improving the convenience of users to understand the meaning of the image.
Smart Images

Figure CN113989589B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an image text detection method, device, storage medium and electronic device. Background Art
[0002] Image text can reflect the meaning and content of an image. Scene text detection is of great value for image understanding and retrieval. The scene text process is mainly divided into two parts: text detection and text recognition. Text detection is to locate the detailed position of the text area in the image, and text recognition is to identify what kind of characters or text are in the area. Text detection is the first step in scene text processing and is crucial to the accuracy of text recognition. In recent years, due to the successful application of natural scene text detection in the Internet industry, scene text detection has become a research hotspot in autonomous driving, scene understanding and product search. Summary of the invention
[0003] In order to overcome the problems existing in the related art, the present disclosure provides an image text detection method, device, storage medium and electronic device.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for detecting text in an image, comprising:
[0005] Extract text area from the image to be detected;
[0006] For each unit text area in the text area, performing text recognition on the unit text area to obtain a recognized text corresponding to the unit text area;
[0007] Determine the writing order of characters in the unit character area according to the recognized text corresponding to the unit character area;
[0008] According to the writing order of the characters, the target recognition text corresponding to the unit character area is determined.
[0009] In some embodiments, determining the writing order of characters in the unit character region according to the recognized text corresponding to the unit character region includes:
[0010] Mask any character in the recognized text corresponding to the unit text area to obtain a masked text;
[0011] Predicting mask characters in the mask text to obtain target predicted characters;
[0012] The writing order of characters in the unit character area is determined according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area.
[0013] In some embodiments, predicting the mask character in the mask text to obtain a target predicted character includes:
[0014] The masked text is input into a pre-trained first language model to predict the masked characters in the masked text to obtain the target predicted characters.
[0015] In some embodiments, determining the writing order of characters in the unit character region according to the target predicted character and the masked characters in the recognized text corresponding to the unit character region includes:
[0016] Calculating a first glyph similarity between the target predicted character and a masked character in the recognized text corresponding to the unit text area;
[0017] When the first glyph similarity is greater than a first preset similarity threshold, determining that the writing order of the characters in the unit character area is a forward sorting, wherein the forward sorting is from top to bottom or from left to right;
[0018] When the first glyph similarity is less than or equal to the first preset similarity threshold, it is determined that the writing order of the characters in the unit character area is reverse sorting, wherein the reverse sorting is from bottom to top or from right to left.
[0019] In some embodiments, predicting the mask character in the mask text to obtain a target predicted character includes:
[0020] Inputting the masked text into a second language model to predict the masked characters in the masked text to obtain a first predicted character;
[0021] Inputting the masked text into a third language model to predict the masked characters in the masked text to obtain a second predicted character;
[0022] The first predicted character and the second predicted character are determined as the target predicted character.
[0023] In some embodiments, determining the writing order of characters in the unit character region according to the target predicted character and the masked characters in the recognized text corresponding to the unit character region includes:
[0024] Calculating a second glyph similarity between the first predicted character and a masked character in the recognized text corresponding to the unit character area, and a third glyph similarity between the second predicted character and the masked character in the recognized text corresponding to the unit character area;
[0025] When the second glyph similarity and the third glyph similarity are both greater than a second preset similarity threshold, determining that the writing order of the characters in the unit character area is a forward sorting, wherein the forward sorting is from top to bottom or from left to right;
[0026] When the second glyph similarity and the third glyph similarity are both smaller than a third preset similarity threshold, it is determined that the writing order of the characters in the unit character area is reverse sorting, wherein the reverse sorting is from bottom to top or from right to left, and wherein the third preset similarity threshold is smaller than the second preset similarity threshold.
[0027] In some embodiments, determining the writing order of characters in the unit character area according to the target predicted character and the masked characters in the recognized text corresponding to the unit character area further includes:
[0028] When the second glyph similarity and the third glyph similarity are not both greater than the second preset similarity threshold, and the second glyph similarity and the third glyph similarity are not both less than the second preset similarity, return to the step of masking any character in the recognized text corresponding to the unit text area.
[0029] In some embodiments, performing text recognition on the unit text area to obtain the recognized text corresponding to the unit text area includes:
[0030] Perform rotation correction on the unit text area;
[0031] Text recognition is performed on the unit text area obtained after rotation correction to obtain the recognized text corresponding to the unit text area.
[0032] In some embodiments, performing rotation correction on the unit text area includes:
[0033] Determine the tilt angle of the unit text area;
[0034] Determining whether the tilt angle is zero;
[0035] When the tilt angle is non-zero, the unit text area is rotated according to the tilt angle so that the tilt angle of the unit text area obtained after the rotation is zero;
[0036] Determine whether the unit text area obtained after rotation is in an upside-down state;
[0037] When the unit character region obtained after rotation is in the upside-down state, the unit character region obtained after rotation is rotated 180 degrees.
[0038] In some embodiments, the step of determining whether the unit text area obtained after the rotation is in an upside-down state includes:
[0039] Perform text recognition on the unit text area obtained after rotation;
[0040] When the recognition result is empty, it is determined that the unit character area obtained after the rotation is in an upside-down state.
[0041] In some embodiments, the performing rotation correction on the unit text area further includes:
[0042] When the tilt angle is zero, determining whether the unit text area is in an upside-down state;
[0043] When the unit character area is in the upside-down state, the unit character area is rotated 180 degrees.
[0044] According to a second aspect of an embodiment of the present disclosure, there is provided an image text detection device, comprising:
[0045] An extraction module is configured to extract a text area from an image to be detected;
[0046] A recognition module is configured to perform text recognition on each unit text area in the text area extracted by the extraction module to obtain a recognition text corresponding to the unit text area;
[0047] A first determination module is configured to determine the writing order of characters in the unit character area according to the recognition text corresponding to the unit character area obtained by the recognition module;
[0048] The second determination module is configured to determine the target recognition text corresponding to the unit text area according to the text writing order determined by the first determination module.
[0049] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the image text detection method provided in the first aspect of the present disclosure are implemented.
[0050] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0051] processor;
[0052] a memory for storing processor-executable instructions;
[0053] Wherein, the processor is configured to: execute the image text detection method provided in the first aspect of the present disclosure.
[0054] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, extract the text area from the image to be detected; then, for each unit text area in the text area, perform text recognition on the unit text area to obtain the recognized text corresponding to the unit text area; next, determine the writing order of the text in the unit text area according to the recognized text corresponding to the unit text area; finally, determine the target recognized text corresponding to the unit text area according to the writing order of the text. In this way, the target recognized text corresponding to the unit text area can be determined according to the writing order of the text in the detected unit text area, so that the sequential text that meets the reading comprehension of ordinary people can be output, breaking through the problem that the traditional detection only mechanically outputs the recognized text results without directionality, thereby making it easier for users to quickly understand the meaning of the image.
[0055] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0057] Figure 1 The figure is a flow chart of a method for detecting text in an image according to an exemplary embodiment of the present disclosure.
[0058] Figure 2A and 2E It is a schematic diagram showing a method of performing rotation correction on a unit text area according to an exemplary embodiment of the present disclosure.
[0059] Figure 2B and 2F It is a schematic diagram showing a method of performing rotation correction on a unit text area according to another exemplary embodiment of the present disclosure.
[0060] Figure 2C and 2G It is a schematic diagram showing a method of performing rotation correction on a unit text area according to another exemplary embodiment of the present disclosure.
[0061] Figure 2D and 2H It is a schematic diagram showing a method of performing rotation correction on a unit text area according to another exemplary embodiment of the present disclosure.
[0062] Figure 3 The present invention is a flowchart of a method for determining the writing order of characters in a unit character area according to a recognized text corresponding to the unit character area according to an exemplary embodiment of the present disclosure.
[0063] Figure 4 The figure is a block diagram of an image text detection device according to an exemplary embodiment of the present disclosure.
[0064] Figure 5 The figure is a block diagram of an image text detection device according to an exemplary embodiment of the present disclosure.
[0065] Figure 6 The figure is a block diagram of an image text detection device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0066] As discussed in the background technology, scene text detection has become a research hotspot for autonomous driving, scene understanding, and product search. However, at this stage, after the text recognition is performed on the image, the text recognition results are simply output mechanically without directionality, that is, the output content is consistent with the order in the original image. In this way, when the text in the image is sorted from right to left, the output document may not conform to the order of ordinary people's reading comprehension, which is not convenient for users to quickly understand the meaning of the image.
[0067] In view of this, the present disclosure provides an image text detection method, device, storage medium and electronic device.
[0068] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0069] Figure 1 This is a flow chart of an image text detection method according to an exemplary embodiment, which is applied to an electronic device, which may be a mobile terminal, such as a mobile phone, a wearable device, a tablet computer, a personal notebook, etc. It may also be a server, such as a local server, a cloud server, etc.
[0070] like Figure 1 As shown, the method may include the following S101 to S104.
[0071] In S101, a text area is extracted from an image to be detected.
[0072] In the present disclosure, a target detection model may be used to extract text areas from an image to be detected.
[0073] Exemplarily, the object detection model can be a real-time scene text detection model with differentiable binarization (DBNet), a connectionist text proposal network (CTPN), or other models.
[0074] In S102, for each unit text region in the text region, text recognition is performed on the unit text region to obtain the recognized text corresponding to the unit text region.
[0075] In the present disclosure, the unit text region can be a text line or a text string.
[0076] In S103, according to the recognized text corresponding to the unit text region, the writing order of the text in the unit text region is determined.
[0077] In the present disclosure, the writing order of the text includes forward sorting and reverse sorting. Among them, forward sorting is from top to bottom or from left to right, and reverse sorting is from bottom to top or from right to left.
[0078] Specifically, when the unit text region is a text line, the forward sorting is from left to right, and the reverse sorting is from right to left; when the unit text region is a text string, the forward sorting is from top to bottom, and the reverse sorting is from bottom to top.
[0079] In S104, according to the writing order of the text, the target recognized text corresponding to the unit text region is determined.
[0080] Specifically, in the case where the writing order of the text is forward sorting, it indicates that the text order in the current recognized text conforms to the direction that is in line with the normal reading and understanding of ordinary people. At this time, the recognized text corresponding to the unit text region is directly determined as the target recognized text corresponding to the unit text region; in the case where the writing order of the text is reverse sorting, it indicates that the text order in the current recognized text does not conform to the direction that is in line with the normal reading and understanding of ordinary people. At this time, the reverse sequence of the recognized text corresponding to the unit text region is determined as the target recognized text corresponding to the unit text region.
[0081] Exemplarily, the writing order of the text is from right to left, that is, reverse sorting. The recognized text corresponding to the unit text region is "Take me home, please". Among them, the reverse sequence of "Take me home, please" is "Please take me home", and at this time, the target recognized text corresponding to the unit text region is "Please take me home".
[0082] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, extract the text area from the image to be detected; then, for each unit text area in the text area, perform text recognition on the unit text area to obtain the recognized text corresponding to the unit text area; next, determine the writing order of the text in the unit text area according to the recognized text corresponding to the unit text area; finally, determine the target recognized text corresponding to the unit text area according to the writing order of the text. In this way, the target recognized text corresponding to the unit text area can be determined according to the writing order of the text in the detected unit text area, so that the sequential text that conforms to the reading comprehension of ordinary people can be output, breaking through the problem that the traditional detection only mechanically outputs the recognized text results without directionality, thereby making it easier for users to quickly understand the meaning of the image.
[0083] The following is a detailed description of the specific implementation method of performing text recognition on the unit text area in the above S102 to obtain the recognized text corresponding to the unit text area. Specifically, it can be achieved by the following steps (1) and (2):
[0084] (1) Perform rotation correction on the unit text area.
[0085] In the present disclosure, the unit text area in the image to be detected may be tilted or inverted. In order to improve the efficiency and accuracy of subsequent text recognition, before performing text recognition on the unit text area, the unit text area is first rotationally corrected so that the unit text area obtained after the rotation correction is horizontal and the font is upright (that is, not inverted).
[0086] (2) Performing text recognition on the unit text area obtained after rotation correction to obtain the recognized text corresponding to the unit text area.
[0087] For example, optical character recognition (OCR) may be used to perform text recognition on the unit text area obtained after rotation correction to obtain a recognized text corresponding to the unit text area.
[0088] The following is a detailed description of the specific implementation method of rotating and correcting the unit text area in the above step (1). Specifically, it can be achieved by following the steps 1) to 7):
[0089] 1) Determine the tilt angle of the unit text area.
[0090] In the present disclosure, when the unit character area is a character row, the inclination angle of the unit character area can be the angle between the main direction of the unit character area (i.e., the direction parallel to the length of the unit character area) and the horizontal rightward direction; when the unit character area is a character column, the inclination angle of the unit character area can be the angle between the main direction of the unit character area (i.e., the direction parallel to the height of the unit character area) and the vertical upward direction.
[0091] For example, Figure 2A As shown, the unit text area is a text line, and its main direction is N→M. At this time, the inclination angle of the unit text area is ∠a.
[0092] For example, Figure 2B As shown, the unit text area is a text line, and its main direction is U→P. At this time, the inclination angle of the unit text area is ∠b.
[0093] For example, Figure 2C As shown, the unit character area is a character string, and its main direction is T→S. At this time, the inclination angle of the unit character area is ∠c.
[0094] For example, Figure 2D As shown, the unit character area is a character string, and its main direction is V→W. At this time, the inclination angle of the unit character area is ∠d.
[0095] 2) Determine whether the tilt angle is zero.
[0096] In the present disclosure, when the inclination angle of the unit text area is zero, it indicates that the unit text area is not tilted, but it may be upside down. Therefore, it is necessary to determine whether the unit text area is in an upside down state, that is, execute the following step 6); when the inclination angle of the unit text area is non-zero, it indicates that the unit text area is tilted. At this time, the unit text area can be rotated according to the inclination angle so that the inclination angle of the unit text area obtained after rotation is zero, that is, execute the following step 3), and then, determine whether the unit text area obtained after rotation is in an upside down state, that is, execute the following step 4).
[0097] 3) Rotate the unit character area according to the tilt angle so that the tilt angle of the unit character area obtained after rotation is zero.
[0098] Specifically, the unit character area may be rotated clockwise by the above-mentioned inclination angle, so that the inclination angle of the unit character area obtained after the rotation is zero.
[0099] For example, Figure 2A As shown in , the inclination angle of the unit text area is ∠a. Rotating it clockwise by ∠a gives Figure 2EThe unit text area shown in (i.e., the unit text area obtained after rotation).
[0100] For example, Figure 2B As shown in , the inclination angle of the unit text area is ∠b. Rotating it clockwise by ∠b gives Figure 2F The unit character area obtained after rotation as shown in (ie, the unit character area obtained after rotation).
[0101] For example, Figure 2C As shown in , the inclination angle of the unit text area is ∠c. Rotating it clockwise by ∠c gives Figure 2G The unit character area obtained after rotation as shown in (ie the unit character area obtained after rotation).
[0102] For example, Figure 2D As shown in , the inclination angle of the unit text area is ∠d. Rotating it clockwise by ∠d, we get Figure 2H The unit character area obtained after rotation as shown in (ie, the unit character area obtained after rotation).
[0103] 4) Determine whether the unit text area obtained after rotation is in an upside-down state.
[0104] When the unit text area obtained after rotation is in an upside-down state (such as Figure 2E , Figure 2H ), perform the following step 5); if the unit text area obtained after rotation is not in an upside-down state (as shown in Figure 2F , Figure 2G ), indicating that the unit text area rotation correction is completed.
[0105] 5) Rotate the unit text area obtained after rotation by 180 degrees.
[0106] 6) Determine whether the unit text area is in an upside-down state.
[0107] When the unit text region is in an upside-down state, perform the following step 7); when the unit text region is not in an upside-down state, it indicates that the unit text region does not need to be rotated and corrected.
[0108] 7) Rotate the unit text area 180 degrees.
[0109] In the above implementation, by detecting the tilt angle and upside-down state of the unit text area, a rotation correction operation of the unit text area at any angle can be implemented, thereby achieving accurate recognition of text in images at any angle.
[0110] The following details the specific implementation for determining whether the unit text region obtained after rotation is in an upside-down state in step 4) above. Specifically, text recognition can be first performed on the unit text region obtained after rotation; in the case where the recognition result is empty, it is determined that the unit text region obtained after rotation is in an upside-down state; in the case where the recognition result is non-empty, it is determined that the unit text region obtained after rotation is not in an upside-down state.
[0111] In addition, step 6) above can be carried out in a manner similar to that for determining whether the unit text region obtained after rotation is in an upside-down state in step 4) above, and will not be elaborated herein in the present disclosure.
[0112] The following details the specific implementation for determining the writing order of the text in the unit text region according to the recognized text corresponding to the unit text region in S103 above. Specifically, it can be achieved through Figure 3 S1031 to S1033 shown in
[0113] In S1031, any character in the recognized text corresponding to the unit text region is masked to obtain a masked text.
[0114] In S1032, the masked character in the masked text is predicted to obtain a target predicted character.
[0115] In S1033, according to the target predicted character and the masked character in the recognized text corresponding to the unit text region, the writing order of the text in the unit text region is determined.
[0116] Exemplarily, the recognized text corresponding to the unit text region is "I'll take you home", and after masking any character therein, the masked text "I'll * you home" is obtained, that is, the masked character is "take".
[0117] The following details the specific implementation for predicting the masked character in the masked text in S1032 above to obtain a target predicted character.
[0118] In one implementation, a pre-trained first language model can be used to predict the masked character in the masked text to obtain a target predicted character. Specifically, the masked text can be input into the first language model to predict the masked character in the masked text, and the predicted character corresponding to the masked character in the masked text, that is, the target predicted character, can be obtained.
[0119] Among them, the first language model can be, for example, a GPT3 (Generative Pre-Training, GPT3) model, a bidirectional encoder representation from transformers (Bidirectional Encoder Representation from Transformers, BERT), etc.
[0120] At this time, the above S1033 can determine the writing order of the characters in the unit character area according to the target predicted character and the masked characters in the recognized text corresponding to the unit character area in the following manner:
[0121] First, the first glyph similarity between the target predicted character and the masked character in the recognized text corresponding to the unit text area is calculated; when the first glyph similarity is greater than a first preset similarity threshold (for example, 0.8), the writing order of the characters in the unit text area is determined to be forward sorted; when the first glyph similarity is less than or equal to the first preset similarity threshold, the writing order of the characters in the unit text area is determined to be reverse sorted.
[0122] In the present disclosure, the first glyph similarity between the target predicted character and the masked character in the recognized text corresponding to the unit text area may be calculated by using an edit distance algorithm, a longest common substring algorithm, or the cosine theorem.
[0123] In another embodiment, the pre-trained second language model and third language model can be used to predict the masked characters in the masked text, respectively, to obtain the first predicted character and the second predicted character as the target predicted character. Specifically, the masked text can be input into the second language model to predict the masked characters in the masked text, to obtain the first predicted character corresponding to the masked characters in the masked text, and at the same time, the masked text can be input into the third language model to predict the masked characters in the masked text, to obtain the second predicted character corresponding to the masked characters in the masked text; then, the first predicted character and the second predicted character are determined as the target predicted characters.
[0124] For example, the second language model may be a GPT3 model, and the third language model may be a BERT model.
[0125] At this time, the above S1033 can determine the writing order of the characters in the unit character area according to the target predicted character and the masked characters in the recognized text corresponding to the unit character area in the following manner:
[0126] First, calculate the second glyph similarity between the first predicted character and the masked character in the recognized text corresponding to the unit text area, and the third glyph similarity between the second predicted character and the masked character in the recognized text corresponding to the unit text area; when the second glyph similarity and the third glyph similarity are both greater than the second preset similarity threshold (for example, 0.7), determine that the writing order of the characters in the unit text area is forward sorting; when the second glyph similarity and the third glyph similarity are both less than the third preset similarity threshold (for example, 0.3), determine that the writing order of the characters in the unit text area is reverse sorting, wherein the third preset similarity threshold is less than the second preset similarity threshold. When the second glyph similarity and the third glyph similarity are not both greater than the second preset similarity threshold, and the second glyph similarity and the third glyph similarity are not both less than the second preset similarity, re-mask any character in the recognized text corresponding to the above unit text area to re-determine the writing order of the characters, that is, return to the above S1031.
[0127] In the present disclosure, the second glyph similarity between the first predicted character and the masked character in the recognized text corresponding to the unit text area and the third glyph similarity between the second predicted character and the masked character in the recognized text corresponding to the unit text area can be calculated by using an edit distance algorithm, a longest common substring algorithm or a cosine theorem.
[0128] In the above implementation, the order of writing characters is determined by using the prediction results of the two language models and the glyph similarity between the masked characters, which can effectively ensure the accuracy of the order of writing characters.
[0129] Figure 4 FIG. 1 is a block diagram of an image text detection device according to an exemplary embodiment. Figure 4 As shown, the device 400 includes:
[0130] An extraction module 401 is configured to extract a text area from an image to be detected;
[0131] The recognition module 402 is configured to perform text recognition on each unit text area in the text area extracted by the extraction module 401 to obtain a recognition text corresponding to the unit text area;
[0132] A first determination module 403 is configured to determine the writing order of characters in the unit character region according to the recognized text corresponding to the unit character region obtained by the recognition module 402;
[0133] The second determination module 404 is configured to determine the target recognition text corresponding to the unit character area according to the character writing order determined by the first determination module 403 .
[0134] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: first, extract the text area from the image to be detected; then, for each unit text area in the text area, perform text recognition on the unit text area to obtain the recognition text corresponding to the unit text area; next, determine the writing order of the text in the unit text area according to the recognition text corresponding to the unit text area; finally, determine the target recognition text corresponding to the unit text area according to the writing order of the text. In this way, the recognition text corresponding to the unit text area can be output according to the writing order of the text in the detected unit text area, so that the sequential copy that conforms to the reading comprehension of ordinary people can be output, breaking through the problem that the traditional detection only mechanically outputs the recognition text result without directionality, thereby making it easier for users to quickly understand the meaning of the image.
[0135] In some embodiments, the first determining module 403 includes:
[0136] The masking submodule is configured to mask any character in the recognized text corresponding to the unit text area to obtain a masked text;
[0137] A prediction submodule, configured to predict mask characters in the mask text to obtain target predicted characters;
[0138] The first determination submodule is configured to determine the writing order of characters in the unit character area according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area.
[0139] In some embodiments, the prediction submodule is configured to input the mask text into a pre-trained first language model to predict the masked characters in the mask text to obtain the target predicted characters.
[0140] In some embodiments, the first determining submodule includes:
[0141] A first calculation submodule is configured to calculate a first glyph similarity between the target predicted character and a masked character in the recognized text corresponding to the unit text area;
[0142] A second determination submodule is configured to determine that the writing order of the characters in the unit character area is a forward order when the first glyph similarity is greater than a first preset similarity threshold, wherein the forward order is from top to bottom or from left to right;
[0143] The third determination submodule is configured to determine that the writing order of the characters in the unit character area is reverse sorting when the first glyph similarity is less than or equal to the first preset similarity threshold, wherein the reverse sorting is from bottom to top or from right to left.
[0144] In some embodiments, the prediction submodule comprises:
[0145] A first input submodule is configured to input the masked text into a second language model to predict masked characters in the masked text to obtain a first predicted character;
[0146] A second input submodule is configured to input the masked text into a third language model to predict the masked characters in the masked text to obtain a second predicted character;
[0147] The predicted character determination submodule is configured to determine the first predicted character and the second predicted character as the target predicted character.
[0148] In some embodiments, the first determining submodule includes:
[0149] A second calculation submodule is configured to calculate a second glyph similarity between the first predicted character and the masked character in the recognized text corresponding to the unit text area, and a third glyph similarity between the second predicted character and the masked character in the recognized text corresponding to the unit text area;
[0150] A fourth determination submodule is configured to determine that the writing order of the characters in the unit character area is a forward sorting when both the second character shape similarity and the third character shape similarity are greater than a second preset similarity threshold, wherein the forward sorting is from top to bottom or from left to right;
[0151] The fifth determination submodule is configured to determine that the writing order of the characters in the unit character area is reverse sorting when the second glyph similarity and the third glyph similarity are both less than a third preset similarity threshold, wherein the reverse sorting is from bottom to top or from right to left, and wherein the third preset similarity threshold is less than the second preset similarity threshold.
[0152] In some embodiments, the first determining submodule further includes:
[0153] The trigger submodule is configured to trigger the mask submodule to mask any character in the recognized text corresponding to the unit text area when the second glyph similarity and the third glyph similarity are not both greater than the second preset similarity threshold and the second glyph similarity and the third glyph similarity are not both less than the second preset similarity.
[0154] In some embodiments, the identification module 402 includes:
[0155] A correction submodule is configured to perform rotation correction on the unit text area;
[0156] The first recognition submodule is configured to perform text recognition on the unit text area obtained after rotation correction to obtain the recognized text corresponding to the unit text area.
[0157] In some embodiments, the correction submodule includes:
[0158] A sixth determination submodule is configured to determine a tilt angle of the unit text area;
[0159] A first judging submodule is configured to judge whether the tilt angle is zero;
[0160] a rotation submodule, configured to rotate the unit text area according to the tilt angle when the tilt angle is non-zero, so that the tilt angle of the unit text area obtained after the rotation is zero;
[0161] The second judgment submodule is configured to judge whether the unit text area obtained after rotation is in an upside-down state;
[0162] The rotation submodule is further configured to rotate the unit character area obtained after the rotation by 180 degrees when the unit character area obtained after the rotation is in the upside-down state.
[0163] In some embodiments, the second determination submodule includes:
[0164] A second recognition submodule is configured to perform text recognition on the unit text area obtained after rotation;
[0165] The seventh determination submodule is configured to determine that the unit character area obtained after the rotation is in an upside-down state when the recognition result is empty.
[0166] In some embodiments, the correction submodule further includes:
[0167] A third judgment submodule is configured to judge whether the unit text area is in an upside-down state when the tilt angle is zero;
[0168] The rotation submodule is further configured to rotate the unit character area by 180 degrees when the unit character area is in the upside-down state.
[0169] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0170] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the image text detection method provided by the present disclosure are implemented.
[0171] The present disclosure also provides an electronic device, comprising:
[0172] processor;
[0173] a memory for storing processor-executable instructions;
[0174] Wherein, the processor is configured to: execute the above-mentioned image text detection method provided by the present disclosure.
[0175] Figure 5 8 is a block diagram of an image text detection device 800 according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0176] Reference Figure 5 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0177] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned image text detection method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0178] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0179] The power component 806 provides power to the various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0180] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0181] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0182] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0183] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0184] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0185] In an exemplary embodiment, the device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above-mentioned image text detection method.
[0186] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of the device 800 to complete the above-mentioned image text detection method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0187] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned image text detection method when executed by the programmable device.
[0188] Figure 6 1 is a block diagram of an image text detection device 1900 according to an exemplary embodiment. For example, the device 1900 may be provided as a server. Figure 6 The device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-mentioned image text detection method.
[0189] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output (I / O) interface 1958. The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or similar.
[0190] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0191] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for detecting text in an image. It is characterized in that include: Extract text area from the image to be detected; For each unit text area in the text area, performing text recognition on the unit text area to obtain a recognized text corresponding to the unit text area; Mask any character in the recognized text corresponding to the unit text area to obtain a masked text; Predicting mask characters in the mask text to obtain target predicted characters; Determine the writing order of characters in the unit character area according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area; According to the writing order of the characters, the target recognition text corresponding to the unit character area is determined.
2. The method according to claim 1, It is characterized in that The predicting the mask character in the mask text to obtain a target predicted character includes: The masked text is input into a pre-trained first language model to predict the masked characters in the masked text to obtain the target predicted characters.
3. The method according to claim 1, It is characterized in that The step of determining the writing order of characters in the unit character area according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area includes: Calculating a first glyph similarity between the target predicted character and a masked character in the recognized text corresponding to the unit text area; When the first glyph similarity is greater than a first preset similarity threshold, determining that the writing order of the characters in the unit character area is a forward sorting, wherein the forward sorting is from top to bottom or from left to right; When the first glyph similarity is less than or equal to the first preset similarity threshold, it is determined that the writing order of the characters in the unit character area is reverse sorting, wherein the reverse sorting is from bottom to top or from right to left.
4. The method according to claim 1, It is characterized in that The predicting the mask character in the mask text to obtain a target predicted character includes: Inputting the masked text into a second language model to predict the masked characters in the masked text to obtain a first predicted character; Inputting the masked text into a third language model to predict the masked characters in the masked text to obtain a second predicted character; The first predicted character and the second predicted character are determined as the target predicted character.
5. The method according to claim 4, It is characterized in that The step of determining the writing order of characters in the unit character area according to the target predicted character and the masked characters in the recognized text corresponding to the unit character area includes: Calculating a second glyph similarity between the first predicted character and a masked character in the recognized text corresponding to the unit character area, and a third glyph similarity between the second predicted character and the masked character in the recognized text corresponding to the unit character area; When the second glyph similarity and the third glyph similarity are both greater than a second preset similarity threshold, determining that the writing order of the characters in the unit character area is a forward sorting, wherein the forward sorting is from top to bottom or from left to right; When the second glyph similarity and the third glyph similarity are both smaller than a third preset similarity threshold, it is determined that the writing order of the characters in the unit character area is reverse sorting, wherein the reverse sorting is from bottom to top or from right to left, and wherein the third preset similarity threshold is smaller than the second preset similarity threshold.
6. The method according to claim 5, It is characterized in that The step of determining the writing order of characters in the unit character area according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area further includes: When the second glyph similarity and the third glyph similarity are not both greater than the second preset similarity threshold, and the second glyph similarity and the third glyph similarity are not both less than the second preset similarity, return to the step of masking any character in the recognized text corresponding to the unit text area.
7. The method according to any one of claims 1 to 6, It is characterized in that The performing text recognition on the unit text area to obtain the recognized text corresponding to the unit text area includes: Perform rotation correction on the unit text area; Text recognition is performed on the unit text area obtained after rotation correction to obtain the recognized text corresponding to the unit text area.
8. The method according to claim 7, It is characterized in that The rotation correction of the unit text area includes: Determine the tilt angle of the unit text area; Determining whether the tilt angle is zero; When the tilt angle is non-zero, the unit text area is rotated according to the tilt angle so that the tilt angle of the unit text area obtained after the rotation is zero; Determine whether the unit text area obtained after rotation is in an upside-down state; When the unit character region obtained after rotation is in the upside-down state, the unit character region obtained after rotation is rotated 180 degrees.
9. The method according to claim 8, It is characterized in that The step of determining whether the unit text area obtained after the rotation is in an upside-down state includes: Perform text recognition on the unit text area obtained after rotation; When the recognition result is empty, it is determined that the unit character area obtained after the rotation is in an upside-down state.
10. The method according to claim 8, It is characterized in that The rotation correction of the unit text area also includes: When the tilt angle is zero, determining whether the unit text area is in an upside-down state; When the unit character area is in the upside-down state, the unit character area is rotated 180 degrees.
11. An image text detection device, It is characterized in that include: An extraction module is configured to extract a text area from an image to be detected; A recognition module is configured to perform text recognition on each unit text area in the text area extracted by the extraction module to obtain a recognition text corresponding to the unit text area; A first determination module is configured to determine the writing order of characters in the unit character area according to the recognition text corresponding to the unit character area obtained by the recognition module; A second determination module is configured to determine the target recognition text corresponding to the unit character area according to the character writing order determined by the first determination module; Wherein, the first determining module includes: The masking submodule is configured to mask any character in the recognized text corresponding to the unit text area to obtain a masked text; A prediction submodule, configured to predict mask characters in the mask text to obtain target predicted characters; The first determination submodule is configured to determine the writing order of characters in the unit character area according to the target predicted characters and the masked characters in the recognized text corresponding to the unit character area.
12. A computer-readable storage medium having computer program instructions stored thereon, It is characterized in that When the program instructions are executed by a processor, the steps of the method described in any one of claims 1 to 10 are implemented.
13. An electronic device, It is characterized in that include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: execute the image text detection method according to any one of claims 1-10.
Citation Information
Patent Citations
Text recognition method and device and electronic equipment
CN112560862A
Preparing a display document for analysis
US20090063965A1
Detecting orientation of textual documents on a live camera feed
US20180314884A1