Text character detection method, device and storage medium

By identifying text images and using the encoding information in the character dictionary to determine the text direction, the problem of inverted text recognition in complex scenarios is solved, and text recognition with high accuracy is achieved.

CN114241184BActive Publication Date: 2025-05-13SF TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010941178.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-09
Publication Date
2025-05-13
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify text rotated inverted 180 degrees in complex scenarios, especially in natural scenes or industrial environments, where the direction of the text is not fixed, resulting in the recognition failure.

Method used

By obtaining the text image to be recognized, character recognition is performed and detection text of multiple characters with an arrangement order is generated. Using the character encoding information in the preset character dictionary, it is determined whether the detection text is forward or inverted text, and outputs its forward text, thereby realizing the recognition of the rotated 180-degree inverted text.

Benefits of technology

The correct recognition of inverted text with 180 degrees is achieved, which expands the usage scenarios of text recognition and improves the recognition accuracy of text in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241184B_ABST
    Figure CN114241184B_ABST
Patent Text Reader

Abstract

The present application discloses a text character detection method, device and storage medium, the text character detection method comprising: obtaining a text image to be recognized; performing character recognition on the text image to be recognized to obtain a detection text of multiple characters in an arrangement order; determining whether the detected text is a forward text or an inverted text according to the character encoding information of the characters in a preset character dictionary, and outputting the forward text of the detected text. The present application can realize the determination of the text direction and output the content recognition result of the forward text corresponding to the detected text, and can complete the recognition regardless of whether the multiple characters after the recognition of the text image to be recognized are forward or inverted, thereby realizing the recognition of bidirectional text, expanding the use scenarios of text recognition, and improving the recognition accuracy of text in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a text character detection method, device and storage medium. Background Art

[0002] At present, text recognition tasks based on deep learning have become relatively mature, but the usual steps of text recognition are to first detect the text area through text detection, then determine the direction of the text, and then perform recognition, or the direction of text recognition is known in advance as the positive direction. For inverted text with the entire text area rotated 180 degrees, it is difficult to obtain correct recognition results if conventional text recognition solutions are used directly.

[0003] Text recognition schemes in natural scenes or industrial environments are much more complicated than the ideal situation where the direction of the text is determined in advance or can be obtained for documents and certificates. For example, the express delivery industry needs to use text recognition technology to identify the text information in the waybill affixed to the package during the transit station. In this scenario, since the placement of the package is not fixed, the location and direction of the waybill are not fixed, resulting in the direction of the image for text recognition is also not fixed. Since each sample is very different, the detection algorithm can only effectively determine the horizontal and vertical directions of the text based on the arrangement rules of the text, and cannot determine the front and back of the text, resulting in the inability to recognize such text normally. Summary of the invention

[0004] The present invention provides a text character detection method, device and storage medium, which can realize the determination of text direction and the content recognition result of the forward text corresponding to the recognized text. Regardless of whether the multiple characters after the text image recognition are forward or inverted, the recognition can be completed, thereby realizing the recognition of bidirectional text, expanding the application scenarios of text recognition, and improving the recognition accuracy of text in complex scenarios.

[0005] On the one hand, the present application provides a text character detection method, the text character detection method comprising:

[0006] Get the text image to be recognized;

[0007] Performing character recognition on the to-be-recognized text image to obtain a detection text having a plurality of characters in an arrangement order;

[0008] Determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text;

[0009] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0010] In some embodiments of the present application, before determining whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in the preset character dictionary and outputting the normal text of the detected text, the method further includes:

[0011] Acquire an initial character set, wherein the initial character set is a character set including a preset number of positive characters, and the characters in the initial character set only include positive characters;

[0012] After inverting the characters in the initial character set, the characters are added to the initial character set to obtain the character set;

[0013] The characters in the character set are encoded to obtain the character dictionary.

[0014] In some implementations of the present application, encoding the characters in the character set to obtain the character dictionary includes:

[0015] Encoding the forward character and the inverted character of the first type of characters in the character set using different encoding information to obtain a first forward character code and a first inverted character code;

[0016] Encoding the forward character and the inverted character of the second type of characters in the character set using different encoding information to obtain a second forward character code and a second inverted character code;

[0017] Encode each character of the third type of characters in the character set and its corresponding character using the same encoding information to obtain a first common character code;

[0018] Encode each character of the fourth type of characters in the character set and its corresponding character using the same encoding information to obtain a second common character code;

[0019] The forward character code includes the first forward character code and the second forward character code, the inverted character code includes the first inverted character code and the second inverted character code, and the common character code includes the first common character code and the second common character code.

[0020] In some embodiments of the present application, determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and outputting the normal text of the detected text includes:

[0021] Determining whether the detected text is a forward text or an inverted text according to character encoding information of characters in a preset character dictionary;

[0022] If the detected text is a positive text, directly output the detected text;

[0023] If the detected text is an inverted text, the detected text is inverted and a result of the inverted text is output.

[0024] In some implementations of the present application, determining whether the detected text is a forward text or an inverted text according to character encoding information of characters in a preset character dictionary includes:

[0025] Taking each character in the detected text as a target character, searching for the character encoding information of the target character in the character encoding information of the character in the preset character dictionary;

[0026] Determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character;

[0027] Counting a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text;

[0028] According to the first quantity value and the second quantity value, it is determined whether the detected text is a normal text or an inverted text.

[0029] In some implementations of the present application, determining whether the detected text is a normal text or an inverted text according to the first quantity value and the second quantity value includes:

[0030] Determine the magnitude of the first quantity value and the second quantity value;

[0031] If the first quantity value is greater than the second quantity value, determining that the detected text is a positive text;

[0032] If the first quantity value is smaller than the second quantity value, it is determined that the detected text is inverted text.

[0033] In some embodiments of the present application, the step of performing character recognition on the text image to be recognized to obtain a detected text having a plurality of characters in an arranged order includes:

[0034] Performing character segmentation on the text image to be recognized to obtain a plurality of character images;

[0035] Character recognition is performed on the multiple character images to obtain a detection text of multiple characters in an arrangement order.

[0036] In some embodiments of the present application, the step of performing character recognition on the text image to be recognized to obtain a detected text having a plurality of characters in an arranged order includes:

[0037] The text image to be recognized is input into a pre-trained text detection model to output a detected text with multiple characters in an arranged order. The text detection model is a DenseNet network model, and the loss function of the DenseNet network model is a weighted temporal connection classification loss function.

[0038] On the other hand, the present application provides a text character detection device, the text character detection device comprising:

[0039] An acquisition unit, used for acquiring a text image to be recognized;

[0040] A recognition unit, used for performing character recognition on the text image to be recognized to obtain a detection text having a plurality of characters in an arrangement order;

[0041] An output unit, used to determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text;

[0042] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0043] In some embodiments of the present application, the device further includes a coding unit, and the coding unit is used to:

[0044] Before determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary and outputting the normal text of the detected text, obtaining an initial character set;

[0045] After inverting the characters in the initial character set, the characters are added to the initial character set to obtain a character set;

[0046] The characters in the character set are encoded to obtain the character dictionary.

[0047] In some implementations of the present application, the encoding unit is specifically used for:

[0048] Encoding the forward character and the inverted character of the first type of characters in the character set using different encoding information to obtain a first forward character code and a first inverted character code;

[0049] Encoding the forward character and the inverted character of the second type of characters in the character set using different encoding information to obtain a second forward character code and a second inverted character code;

[0050] The forward character and the inverted character of the target character in the character set are respectively encoded using the same encoding information to obtain a common character code, wherein the target character is a character whose forward character and inverted character are the same as the character in the character set, or a character whose forward character is the same as the forward character of other characters in the character set;

[0051] The forward character code includes the first forward character code and the second forward character code, and the inverted character code includes the first inverted character code and the second inverted character code.

[0052] In some embodiments of the present application, the output unit is specifically used for:

[0053] Determining whether the detected text is a forward text or an inverted text according to character encoding information of characters in a preset character dictionary;

[0054] If the detected text is a positive text, directly output the detected text;

[0055] If the detected text is an inverted text, the detected text is inverted and a result of the inverted text is output.

[0056] In some embodiments of the present application, the output unit is specifically used for:

[0057] Taking each character in the detected text as a target character, searching for the character encoding information of the target character in the character encoding information;

[0058] Determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character;

[0059] Counting a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text;

[0060] According to the first quantity value and the second quantity value, it is determined whether the detected text is a normal text or an inverted text.

[0061] In some embodiments of the present application, the output unit is specifically used for:

[0062] Determine the magnitude of the first quantity value and the second quantity value;

[0063] If the first quantity value is greater than the second quantity value, determining that the detected text is a positive text;

[0064] If the first quantity value is smaller than the second quantity value, it is determined that the detected text is inverted text.

[0065] In some embodiments of the present application, the identification unit is specifically used for:

[0066] Performing character segmentation on the text image to be recognized to obtain a plurality of character images;

[0067] Character recognition is performed on the multiple character images to obtain a detection text of multiple characters in an arrangement order.

[0068] In some embodiments of the present application, the identification unit is specifically used for:

[0069] The text image to be recognized is input into a pre-trained text detection model to output a detected text with multiple characters in an arranged order. The text detection model is a DenseNet network model, and the loss function of the DenseNet network model is a weighted temporal connection classification loss function.

[0070] On the other hand, the present application also provides a computer device, the computer device comprising:

[0071] one or more processors;

[0072] Memory; and

[0073] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the processor to implement the text character detection method described in any one of the first aspects.

[0074] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is loaded by a processor to execute the steps in the text character detection method described in any one of the first aspects.

[0075] In the present application, a text image to be recognized is obtained; character recognition is performed on the text image to be recognized to obtain a detection text of multiple characters with an arrangement order; according to the character encoding information of the characters in the preset character dictionary, it is determined whether the detection text is a normal text or an inverted text, and the normal text of the detection text is output. The present application pre-encodes the characters in the character dictionary as normal character encoding, inverted character encoding and common character encoding, and then based on the character encoding information of the characters in the character dictionary, character direction recognition is performed on multiple characters after the text image to be recognized is determined to determine the normal text of the multiple characters, so that the determination of the direction of the detection text and the output of the content recognition result of the normal text corresponding to the detection text can be realized, regardless of whether the multiple characters after the text image to be recognized are normal or inverted, the recognition can be completed, thereby realizing the recognition of bidirectional text, expanding the use scenarios of text recognition, and improving the recognition accuracy of text in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0077] Figure 1 is a schematic diagram of a scenario of a text character detection system provided by an embodiment of the present invention;

[0078] Figure 2 is a schematic flow chart of an embodiment of a text character detection method provided in an embodiment of the present invention;

[0079] Figure 3 is a schematic flow chart of an embodiment of step 203 in an embodiment of the present invention;

[0080] Figure 4 is a schematic flow chart of an embodiment of step 301 in an embodiment of the present invention;

[0081] Figure 5 is a schematic diagram of the structure of an embodiment of a text character detection device provided in an embodiment of the present invention;

[0082] Figure 6 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0083] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0084] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0085] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any person skilled in the art to implement and use the invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.

[0086] The embodiments of the present invention provide a text character detection method, device and storage medium, which are described in detail below.

[0087] See also Figure 1 , Figure 1 Schematic diagram of a text character detection system provided by an embodiment of the present invention. The text character detection system may include a computer device 100, in which a text character detection device is integrated. Figure 1 Computer equipment in.

[0088] In the embodiment of the present invention, the computer device 100 is mainly used to obtain a text image to be recognized; perform character recognition on the text image to be recognized to obtain a detected text of multiple characters with an arrangement order; determine whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in a preset character dictionary, and output the normal text of the detected text; wherein the character encoding information of the characters in the character dictionary includes character encoding information of first type characters, second type characters, third type characters and fourth type characters predefined in a preset character set, the first type characters are characters whose normal characters and inverted characters are different, the second type characters are characters whose inverted characters are similar to the normal characters of other characters in the character set or characters whose normal characters are similar to the inverted characters of other characters in the character dictionary, and the third type characters are characters whose normal characters are different from the inverted characters of other characters in the character set. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes information of pre-set forward character encoding, inverted character encoding and common character encoding, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0089] In the embodiment of the present invention, the computer device 100 may be an independent server, or a server network or server cluster composed of servers. For example, the computer device 100 described in the embodiment of the present invention includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of a plurality of servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0090] It is understandable that the computer device 100 used in the embodiment of the present invention may also be a device including both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. Such a device may include: a cellular or other communication device having a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. The specific computer device 100 may be a desktop terminal or a mobile terminal, for example, the computer device 100 may also be a tablet computer, a laptop computer, etc.

[0091] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or less computer equipment as shown in Figure 1 Only one computer device is shown in the figure. It can be understood that the text character detection system can also include one or more other computer devices, which are not specifically limited here.

[0092] In addition, if Figure 1 As shown, the text character detection system may further include a memory 200 for storing data, such as character data, such as characters in a character dictionary or characters in a detected text.

[0093] It should be noted that Figure 1 The scenario diagram of the text character detection system shown is merely an example. The text character detection system and scenario described in the embodiment of the present invention are intended to more clearly illustrate the technical solution of the embodiment of the present invention, and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. A person of ordinary skill in the art can appreciate that with the evolution of the text character detection system and the emergence of new business scenarios, the technical solution provided by the embodiment of the present invention is equally applicable to similar technical problems.

[0094] First, a text character detection method is provided in an embodiment of the present invention. The execution subject of the text character detection method is a text character detection device, and the text character detection device is applied to a computer device. The text character detection method comprises: obtaining a text image to be recognized; performing character recognition on the text image to be recognized to obtain a detection text of multiple characters with an arrangement order; determining whether the detection text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and outputting the normal text of the detection text; wherein the character encoding information of the characters in the character dictionary comprises character encoding information of a first type of characters, a second type of characters, a third type of characters and a fourth type of characters predefined in a preset character set, the first type of characters being characters whose normal characters and inverted characters are different, and the second type of characters being characters whose inverted characters are similar to normal characters of other characters in the character set or characters whose normal characters and inverted characters are different. the character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters include information of pre-set forward character encoding, inverted character encoding and common character encoding, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0095] like Figure 2 FIG. 2 is a flow chart of an embodiment of a text character detection method according to an embodiment of the present invention. The text character detection method includes steps 201 to 203, which are as follows:

[0096] 201. Obtain a text image to be recognized.

[0097] Among them, the text image to be recognized can be an image of the text to be recognized taken by a shooting device. For example, in the logistics industry, when transferring packages at logistics outlets, text recognition technology is needed to recognize the text information in the waybill affixed to the package. At this time, the text image to be recognized is the waybill image of the package taken by a shooting device in the logistics outlet, where the logistics outlet can be a transfer yard or a logistics collection and delivery outlet.

[0098] 202. Perform character recognition on the to-be-recognized text image to obtain a detected text having a plurality of characters in an arrangement order.

[0099] Among them, the arrangement order is the arrangement order that matches the text image to be recognized. Specifically, it can be the order corresponding to each character area in the text to be recognized, for example, the order of the detection text box corresponding to each character in the text image to be recognized (which can be the detection box of the neural network in subsequent embodiments).

[0100] 203. Determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text.

[0101] Among them, the character dictionary is a collection of character codes. In the embodiment of the present application, the character dictionary may include character codes of character types such as numeric characters, alphabetic characters, and symbol characters. For alphabetic characters, it may also include uppercase and lowercase alphabetic characters, such as A and a. In addition, for each character type, it may include the forward character and inverted character of the character. For example, for the number A, the forward character is "A" and the inverted character is The character dictionary can include "A" and It can be understood that in certain embodiments of the present application, the character dictionary may further include characters of other computer languages ​​or human languages. Since there are many characters in other computer languages ​​or human languages, the situation is more complicated. When the number of characters is too large, it is impossible to define all characters exhaustively. Therefore, the present application is mainly used for character encoding of character types including numeric characters, alphabetic characters, and symbolic characters in the character dictionary, that is, the detection scenario of the present application preferably recognizes text images corresponding to numeric characters, alphabetic characters, and symbolic characters.

[0102] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0103] Among them, since the third type of characters are characters whose forward character and inverted character are the same, and the fourth type of characters are characters whose inverted character and forward characters of other characters in the character set are the same, at this time, each character in the third type of characters and its corresponding character represent the forward character and inverted character of each character in the third type of characters; each character in the fourth type of characters and its corresponding character represent the inverted character of the characters in the fourth type of characters and the forward character of other characters in the character set are the same.

[0104] The present application pre-encodes the characters in the character dictionary into forward character encoding, inverted character encoding and common character encoding, and then based on the character encoding information of the characters in the character dictionary, performs character direction recognition on multiple characters after the text image to be recognized, and determines the forward text of the multiple characters. Therefore, it is possible to determine the direction of the detected text and output the content recognition result of the forward text corresponding to the detected text. Regardless of whether the multiple characters after the text image to be recognized are forward or inverted, the recognition can be completed, thereby realizing the recognition of bidirectional text, expanding the use scenarios of text recognition, and improving the recognition accuracy of text in complex scenarios.

[0105] When using deep learning solutions for text recognition, each character is essentially encoded, and the recognition process is transformed into a character classification problem, with each character corresponding to its encoding category. Therefore, the model does not care whether the text is inverted. The inverted characters and the forward characters are combined to form the character dictionary, and the characters in the character dictionary can be used to train the corresponding text detection model.

[0106] Therefore, in some embodiments of the present application, before determining whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in the preset character dictionary and outputting the normal text of the detected text, the method may also include: obtaining an initial character set; inverting the characters in the initial character set and adding them to the initial character set to obtain a character set; encoding the characters in the character set to obtain the character dictionary. At this time, the character encoding information of the characters in the character dictionary includes information of pre-set normal character encoding, inverted character encoding and common character encoding. Among them, the initial character set is a character set including a preset number of normal characters, and the characters in the initial character set only include normal characters, for example, normal characters of character types such as numbers, letters and symbols, and the initial character set is such as the character dictionary in the prior art, and the characters in the initial character set do not include inverted characters of characters, but only include normal characters.

[0107] When encoding characters in a character dictionary, the simplest way is to treat all characters in the normal text and the characters after the normal text is inverted as different individuals and encode them differently. However, there are actually the following four character occurrence situations between the normal text and the inverted text (hereinafter referred to as "character situations"):

[0108] 1. The first type of characters: the inverted characters and the normal characters of the characters are different;

[0109] 2. The third type of characters: the forward and reverse characters of the characters are the same, such as the forward and reverse characters of "1" are the same;

[0110] 3. The fourth type of characters: the inverted character of a character is the same as the normal character of another character, such as the inverted character of "6" is the same as the normal character of "9";

[0111] 4. The second type of characters: the inverted character of a character is similar to the normal character of other characters, such as the inverted character of "5" is similar to the normal character of "s".

[0112] The above simple encoding method is only valid if all characters only exist in the first case. In actual scenarios, there are four types of characters. Since the forward and reverse characters of the characters are inconsistent, they can be regarded as different individual codes. Therefore, if the above simple encoding method is used, there will be the following two problems:

[0113] (1) In the above-mentioned second and third situations, images with consistent image features use different encodings, which can easily lead to confusion during model training and recognition, directly affecting the performance of the model. When this error occurs, it will directly affect the post-processing decoding. If the direction of the text cannot be determined based on the recognition result, the recognition result cannot be decoded.

[0114] (2) Similar characters in the fourth case are already a difficult point in text recognition, and it is impossible to solve the problem of similar characters being easily misrecognized by simply adjusting the encoding method.

[0115] Therefore, corresponding solutions are designed for the above-mentioned problems.

[0116] First, let's introduce character encoding. Character encoding (English: Character encoding) is also called character set code, which encodes characters in a character set into an object in a specified set (for example: bit pattern, natural number sequence, 8-bit group or electric pulse) so that text can be stored in a computer and transmitted through a communication network. Common examples include encoding the Latin alphabet into Morse code and ASCII code. Among them, ASCII code numbers letters, numbers and other symbols, and represents this integer with 7 bits of binary, usually using an additional extended bit to facilitate storage in 1 byte.

[0117] In response to problem (1), the present application designs a character encoding method, 1. Uniquely encodes two characters in the first type of characters to ensure that the encoding information corresponding to the forward text and inverted text of the characters in the first type of characters are unique. 2. Use the same encoding information for the two characters corresponding to the third type of characters and the corresponding two characters in the fourth type of characters, that is, the two characters in the third type of characters and the two characters in the fourth type of characters use the same encoding information, for example, the forward and inverted "1" use the same encoding method, that is, the same encoding information, for example, in the ASCII decimal encoding rule, the encoding information of 1 is "049", at this time, the forward and inverted "1" both use the same "049" encoding information, the forward "6" and the inverted "9" use the same encoding information, for example, in the ASCII decimal encoding rule, the encoding of the forward "6" is "054", at this time, the forward "6" and the inverted "9" are both encoded with "054". Since the same encoding information is used, their display in the character dictionary is actually one character, for example, the forward "6" and the inverted "9" are the same "054" encoding information.

[0118] Regarding problem (2), adjusting the encoding rules cannot solve the problem of easy errors in similar text recognition. Therefore, this application has formulated a corresponding strategy to solve this problem during model training, and selected simple unique encoding information for the encoding of two characters in the fourth type of characters.

[0119] Specifically, the encoding of the characters in the character set to obtain the character dictionary includes: encoding the forward characters and the inverted characters of the first type of characters in the character set using different encoding information respectively to obtain a first forward character code and a first inverted character code; encoding the forward characters and the inverted characters of the second type of characters in the character set using different encoding information respectively to obtain a second forward character code and a second inverted character code; encoding each character of the third type of characters in the character set and its corresponding characters using the same encoding information to obtain a first common character code; encoding each character of the fourth type of characters in the character set and its corresponding characters using the same encoding information to obtain a second common character code; wherein the forward character code includes the first forward character code and the second forward character code, the inverted character code includes the first inverted character code and the second inverted character code, and the common character code includes the first common character code and the second common character code.

[0120] Among them, the encoding in the embodiment of the present application is divided into three cases, the forward character encoding corresponding to the forward characters in the first type of characters and the second type of characters, the inverted character encoding corresponding to the inverted characters in the first type of characters and the second type of characters, and the common character encoding corresponding to the forward characters and inverted characters in the third type of characters and the fourth type of characters.

[0121] Based on the above principles, in an embodiment of the present application, the character encoding information of the characters in the character dictionary includes character encoding information of a first type of characters, a second type of characters, a third type of characters, and a fourth type of characters pre-defined in a preset character set, the first type of characters being characters whose forward characters and inverted characters are different (the above character case 1), the second type of characters being characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary (the above character case 4), the third type of characters being characters whose forward characters and inverted characters are the same as the characters (the above character case 2), the fourth type of characters being characters whose inverted characters and other characters in the character set are different. Characters that are the same as the forward characters (character situation 3 above); the character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0122] like Figure 3 As shown, in some embodiments of the present application, the step 203 of determining whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in the preset character dictionary, and outputting the normal text of the detected text may include steps 301 to 303, which are specifically as follows:

[0123] 301. Determine whether the detected text is a forward text or an inverted text according to character encoding information of characters in a preset character dictionary.

[0124] 302. If the detected text is a positive text, directly output the detected text.

[0125] 303. If the detected text is an inverted text, perform inversion processing on the detected text and output a result of the inverted processing of the detected text.

[0126] Among them, Figure 4 As shown, in some embodiments of the present application, the step 301 of determining whether the detected text is a forward text or an inverted text according to the character encoding information of the characters in the preset character dictionary may further include steps 401 to 403, which are as follows:

[0127] 401. Determine whether each character in the detected text is a forward character code or an inverted character code according to character code information of the characters in the preset character dictionary.

[0128] Specifically, the method of determining whether each character in the detected text is a forward character code or an inverted character code according to the character code information of the characters in the preset character dictionary includes: taking each character in the detected text as a target character, searching for the character code information of the target character in the character code information of the characters in the preset character dictionary; and determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character.

[0129] For example, the detected text includes the character "A", and through the character encoding information of the character in the dictionary, it can be determined that the character "A" is a forward character encoding.

[0130] 402. Count a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text.

[0131] For example, the detected text includes 5 characters, the number of forward character codes (ie, the first quantity value) is 5, and the number of inverted character codes (ie, the second quantity value) is 1.

[0132] 403. Determine whether the detected text is a normal text or an inverted text according to the first quantity value and the second quantity value.

[0133] Specifically, in some embodiments of the present application, determining whether the detected text is normal text or inverted text based on the first quantity value and the second quantity value may further include: judging the size of the first quantity value and the second quantity value; if the first quantity value is greater than the second quantity value, determining that the detected text is normal text; if the first quantity value is less than the second quantity value, determining that the detected text is inverted text.

[0134] Continuing with the example description in step 402, if the number of normal character codes is 5> the number of inverted character codes is 1, it can be determined that the detected text is normal text. Conversely, if the first number value is less than the second number value, it can be determined that the detected text is inverted text.

[0135] It should be noted that in actual application, there are still very few extreme cases, that is, the first quantity value is equal to the second quantity value. For the equal case, it means that the above method cannot determine whether the detected text at this time is a normal text or an inverted text. At this time, it can be processed according to the actual situation. When the first quantity value is equal to the second quantity, it is determined whether the detected text is a normal text or an inverted text. There are multiple processing methods in the embodiment of the present application:

[0136] (1) Directly discard the detected text without confirming it

[0137] That is, if the first quantity value is equal to the second quantity value, it is determined that the image of the text to be recognized is an image in which it is impossible to determine whether the detected text is positive text or inverted text, and it can be directly discarded without processing, and a reminder message is output to prompt the user.

[0138] (2) Directly determine the detected text according to a fixed text direction

[0139] That is, if the first quantity value is equal to the second quantity value, according to the pre-set fixed text direction (forward or inverted), it is determined whether the detected text is forward text or inverted text. Of course, according to the a priori rule, the inventor believes that when the number of forward characters and the number of reverse characters are equal, the number of forward text cases is larger and the probability of forward text is greater. Therefore, in the embodiment of the present application, preferably, if the first quantity value is equal to the second quantity value, it can be directly determined that the detected text is forward text.

[0140] In the embodiment of the present application, no matter whether the direction of the text in the text image to be recognized is forward or inverted, the expected recognition result is the character recognition result of the order of normal language habits from left to right after the forward text is adjusted. Therefore, for the inverted text, after the text character recognition, it is necessary to make a corresponding judgment on the direction of the recognized text to further determine the output result. After obtaining the detection text including multiple characters, the present application compares the number of characters falling in the forward coding interval (i.e., the number of forward character codes in these characters) and the inverted coding interval (i.e., the number of inverted character codes in these characters) to realize the judgment of the forward and inverted text. Among them, for the characters falling in the common coding interval (i.e., the common character coding), since the judgment on the forward and inverted does not work, it can be directly eliminated and not counted.

[0141] In the embodiment of the present application, on the one hand, the number of normal and inverted characters is determined to determine whether the entire text to be recognized is inverted or normal. At the same time, the text judgment error can be further corrected after the judgment. For example, 5 characters are recognized, 3 of which are normal and 2 are inverted, then the overall text is normal, and the 2 inverted ones are recognition errors that need to be corrected, and finally the normal text of the detected text is output.

[0142] In the embodiment of the present application, the character recognition of the text image to be recognized in step 203 to obtain the detected text with multiple characters in an arranged order can be recognized by a pre-trained neural network model or by an algorithm.

[0143] Specifically, the step 203 of performing character recognition on the text image to be recognized to obtain a detected text with multiple characters in an arranged order may further include: performing character segmentation on the text image to be recognized to obtain multiple character images; performing character recognition on the multiple character images to obtain a detected text with multiple characters in an arranged order. Character segmentation is to segment the area where each character in the text image to be recognized is located. After performing character segmentation on the text image to be recognized, the character image of each character in the text image to be recognized can be obtained.

[0144] Among them, the method of performing character segmentation on the text image to be recognized to obtain multiple character images can adopt the currently existing character image segmentation method, such as first performing grayscale processing, then binarization processing, extracting features after correcting the image, and implementing character segmentation through classifiers such as Support Vector Machine (SVM) and Artificial Neural Network (ANN) to obtain multiple character images. The character recognition of the multiple character images can be performed using the existing OCR algorithm, which is not limited here.

[0145] In some other embodiments of the present application, recognition can be performed directly through a pre-trained neural network model (such as a text detection model), and the character recognition is performed on the text image to be recognized to obtain a detection text with multiple characters in an arranged order, including: inputting the text image to be recognized into a pre-trained text detection model to output a detection text with multiple characters in an arranged order, and the text detection model is a DenseNet network model, and the loss function of the DenseNet network model is a weighted temporal connection classification loss function.

[0146] The preset neural network model may be a convolutional neural network (CNN) model, such as a DenseNet network model, a ResNet network model, a GoogLeNet network model, etc. in the convolutional neural network (CNN) model. Based on the inventor's test, in the embodiment of the present application, the DenseNet network model has better performance with fewer parameters and computational costs, and the preset neural network model is preferably a DenseNet network model.

[0147] Furthermore, when the preset neural network model is a DenseNet network model, its loss function may be a weighted time connection classification loss function (Weight Connectionist temporal classification Loss, Weight CTC Loss) to further improve the character recognition performance.

[0148] For the second type of characters in the character encoding, the recognition results are prone to errors because the characters are too similar. In the embodiment of the present application, a weighted loss is designed to solve the above problem. A greater penalty is given for recognition errors between similar texts. Under the premise of not affecting the performance of the model in recognizing non-similar texts, the recognition ability of the network model on similar texts is improved.

[0149] In the embodiment of the present application, Weight CTC Loss is shown as follows:

[0150]

[0151] Among them, N is the number of characters in the character sample set during text detection model training, K is the number of categories, and y ik is the true label of the ith ik is the predicted probability of the i-th sample in the k-th category (i.e., forward character encoding, inverted character encoding, or common character encoding), α is the weight coefficient, in a specific embodiment of text detection model training, when training the preset neural network model, the character sample and the actual result of the character sample are used for training, for example, the sample is a character of the second type of character, the actual result is already known, if the model output is not this result, it can be judged as a recognition error. When using the sample to train the preset neural network model, when the sample has a recognition error of a non-second type of character, α = 1.0, when the sample has a recognition error in the second type of character, α = 1.1.

[0152] It should be noted that the above-mentioned value of α is only an example. It can be understood that when the sample has character recognition errors of non-second type characters and character recognition errors of second type characters, the value of α can be other values. It only needs to satisfy that the value of α is larger when the sample has character recognition errors of second type characters, and the difference between the values ​​of α in the two cases is within 5-15%.

[0153] Before using the text detection model in the embodiment of the present application, it is necessary to pre-train the text detection model. Specifically, the embodiment of the present application also includes a model training process: taking the characters in the character dictionary as character samples and adding them to the character sample set; training the preset neural network model according to the character sample set to obtain the text detection model. Subsequently, character recognition can be performed based on the text image to be recognized described by the text detection model. In the embodiment of the present application, in order to ensure the subsequent recognition accuracy, the character sample set at least includes the characters in the character dictionary in the embodiment of the present application, so the number of characters in the character sample set can be the total number of characters in the character dictionary in the present application.

[0154] In addition, when training the text detection model in the embodiment of the present application, a fixed word frequency can be used to generate text data in the first type of characters in the above-mentioned character encoding. Since the third type of characters, the second type of characters, and the fourth type of characters account for a smaller proportion of the total number of characters relative to the first type of characters, in the process of data generation, character samples with X times (X is a positive integer) a fixed word frequency are generated for them to solve the problem of too low sample size to accurately and effectively identify them.

[0155] In order to better implement the text character detection method in the embodiment of the present application, based on the text character detection method, the embodiment of the present application also provides a text character detection device, such as Figure 5 As shown, the text character detection device 500 includes an acquisition unit 501, a recognition unit 502 and an output unit 503, which are specifically as follows:

[0156] An acquisition unit 501 is used to acquire a text image to be recognized;

[0157] The recognition unit 502 is used to perform character recognition on the text image to be recognized to obtain a detected text with a plurality of characters in an arranged order;

[0158] An output unit 503, used to determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text;

[0159] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0160] In the present application, an image of a text to be recognized is obtained; character recognition is performed on the image of the text to be recognized to obtain a detection text of multiple characters with an arrangement order; according to the character encoding information of the characters in a preset character dictionary, it is determined whether the detection text is a normal text or an inverted text, and the normal text of the detection text is output; the present application pre-encodes the characters in the character dictionary as normal character encoding, inverted character encoding and common character encoding, and then based on the character encoding information of the characters in the character dictionary, decodes the multiple characters after the image of the text to be recognized is recognized, and determines the normal text of the multiple characters, so that the determination of the text direction and the output of the content recognition result of the normal text corresponding to the detection text can be realized, regardless of whether the multiple characters after the image of the text to be recognized are normal or inverted, the recognition can be completed, thereby realizing the recognition of bidirectional text, expanding the application scenarios of text recognition, and improving the recognition accuracy of text in complex scenarios.

[0161] In some embodiments of the present application, the device further includes a coding unit, and the coding unit is used to:

[0162] Before determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary and outputting the normal text of the detected text, obtaining an initial character set;

[0163] After inverting the characters in the initial character set, the characters are added to the initial character set to obtain a character set;

[0164] Encode the characters in the character set to obtain the character dictionary.

[0165] In some implementations of the present application, the encoding unit is specifically used for:

[0166] Encoding the forward character and the inverted character of the first type of characters in the character set using different encoding information to obtain a first forward character code and a first inverted character code;

[0167] Encoding the forward character and the inverted character of the second type of characters in the character set using different encoding information to obtain a second forward character code and a second inverted character code;

[0168] The forward character and the inverted character of the target character in the character set are respectively encoded using the same encoding information to obtain a common character code, wherein the target character is a character whose forward character and inverted character are the same as the character in the character set, or a character whose forward character is the same as the forward character of other characters in the character set;

[0169] The forward character code includes the first forward character code and the second forward character code, and the inverted character code includes the first inverted character code and the second inverted character code.

[0170] In some implementations of the present application, the output unit 503 is specifically used to:

[0171] Determining whether the detected text is a forward text or an inverted text according to character encoding information of characters in a preset character dictionary;

[0172] If the detected text is a positive text, directly output the detected text;

[0173] If the detected text is an inverted text, the detected text is inverted and a result of the inverted text is output.

[0174] In some implementations of the present application, the output unit 503 is specifically used to:

[0175] Taking each character in the detected text as a target character, searching for the character encoding information of the target character in the character encoding information;

[0176] Determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character;

[0177] Counting a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text;

[0178] According to the first quantity value and the second quantity value, it is determined whether the detected text is a normal text or an inverted text.

[0179] In some implementations of the present application, the output unit 503 is specifically used to:

[0180] Determine the magnitude of the first quantity value and the second quantity value;

[0181] If the first quantity value is greater than the second quantity value, determining that the detected text is a positive text;

[0182] If the first quantity value is smaller than the second quantity value, it is determined that the detected text is inverted text.

[0183] In some implementations of the present application, the identification unit 502 is specifically used for:

[0184] Performing character segmentation on the text image to be recognized to obtain a plurality of character images;

[0185] Character recognition is performed on the multiple character images to obtain a detection text of multiple characters in an arrangement order.

[0186] In some implementations of the present application, the identification unit 502 is specifically used for:

[0187] The text image to be recognized is input into a pre-trained text detection model to output a detected text with multiple characters in an arranged order. The text detection model is a DenseNet network model, and the loss function of the DenseNet network model is a weighted temporal connection classification loss function.

[0188] The embodiment of the present invention further provides a computer device, which integrates any text character detection device provided by the embodiment of the present invention, and the computer device includes:

[0189] one or more processors;

[0190] Memory; and

[0191] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the processor to perform the steps of the text character detection method described in any of the above-mentioned text character detection method embodiments.

[0192] The embodiment of the present invention further provides a computer device, which integrates any text character detection device provided by the embodiment of the present invention. Figure 6As shown, it shows a schematic diagram of the structure of a computer device involved in an embodiment of the present invention, specifically:

[0193] The computer device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will appreciate that Figure 6 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:

[0194] The processor 601 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 602 and calling data stored in the memory 602, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 601.

[0195] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0196] The computer device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 603 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0197] The computer device may further include an input unit 604, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0198] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 601 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602, thereby realizing various functions, as follows:

[0199] Get the text image to be recognized;

[0200] Performing character recognition on the to-be-recognized text image to obtain a detection text having a plurality of characters in an arrangement order;

[0201] Determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text;

[0202] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0203] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0204] To this end, an embodiment of the present invention provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any text character detection method provided in an embodiment of the present invention. For example, the computer program is loaded by a processor to execute the following steps:

[0205] Get the text image to be recognized;

[0206] Performing character recognition on the to-be-recognized text image to obtain a detection text having a plurality of characters in an arrangement order;

[0207] Determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text;

[0208] The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

[0209] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.

[0210] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.

[0211] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0212] The above is a detailed introduction to a text character detection method, device and storage medium provided in an embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A text character detection method, characterized in that: The method comprises: Get the text image to be recognized; Performing character recognition on the to-be-recognized text image to obtain a detection text having a plurality of characters in an arrangement order; Determine whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and output the normal text of the detected text; The step of determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary includes: Taking each character in the detected text as a target character, searching for the character encoding information of the target character in the character encoding information; Determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character; Counting a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text; The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

2. The text character detection method according to claim 1, characterized in that: Before determining whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in the preset character dictionary and outputting the normal text of the detected text, the method further includes: Acquire an initial character set, wherein the initial character set is a character set including a preset number of positive characters, and the characters in the initial character set only include positive characters; After inverting the characters in the initial character set, the characters are added to the initial character set to obtain the character set; The characters in the character set are encoded to obtain the character dictionary.

3. The text character detection method according to claim 2, characterized in that: The encoding of the characters in the character set to obtain the character dictionary includes: Encoding the forward character and the inverted character of the first type of characters in the character set using different encoding information to obtain a first forward character code and a first inverted character code; Encoding the forward character and the inverted character of the second type of characters in the character set using different encoding information to obtain a second forward character code and a second inverted character code; Encode each character of the third type of characters in the character set and its corresponding character using the same encoding information to obtain a first common character code; Encode each character of the fourth type of characters in the character set and its corresponding character using the same encoding information to obtain a second common character code; The forward character code includes the first forward character code and the second forward character code, the inverted character code includes the first inverted character code and the second inverted character code, and the common character code includes the first common character code and the second common character code.

4. The text character detection method according to claim 1, characterized in that: The step of determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary, and outputting the normal text of the detected text comprises: Determining whether the detected text is a forward text or an inverted text according to the character encoding information; If the detected text is a positive text, directly output the detected text; If the detected text is an inverted text, the detected text is inverted and a result of the inverted text is output.

5. The text character detection method according to claim 4, characterized in that: The step of determining whether the detected text is a normal text or an inverted text according to character encoding information of characters in a preset character dictionary includes: Taking each character in the detected text as a target character, searching for the character encoding information of the target character in the character encoding information; Determining whether the target character is a forward character code or an inverted character code according to the character code information of the target character; Counting a first quantity value of the forward character code and a second quantity value of the inverted character code in the detected text; According to the first quantity value and the second quantity value, it is determined whether the detected text is a normal text or an inverted text.

6. The text character detection method according to claim 5, characterized in that: The step of determining whether the detected text is a normal text or an inverted text according to the first quantity value and the second quantity value includes: Determine the magnitude of the first quantity value and the second quantity value; If the first quantity value is greater than the second quantity value, determining that the detected text is a positive text; If the first quantity value is smaller than the second quantity value, it is determined that the detected text is inverted text.

7. The text character detection method according to claim 1, characterized in that: The step of performing character recognition on the to-be-recognized text image to obtain a detected text having a plurality of characters in an arranged order comprises: Performing character segmentation on the text image to be recognized to obtain a plurality of character images; Character recognition is performed on the multiple character images to obtain a detection text of multiple characters in an arrangement order.

8. The text character detection method according to claim 1, characterized in that: The step of performing character recognition on the to-be-recognized text image to obtain a detected text having a plurality of characters in an arranged order comprises: The text image to be recognized is input into a pre-trained text detection model to output a detected text with multiple characters in an arranged order. The text detection model is a DenseNet network model, and the loss function of the DenseNet network model is a weighted temporal connection classification loss function.

9. A text character detection device, characterized in that: The text character detection comprises: An acquisition unit, used for acquiring a text image to be recognized; A recognition unit, used for performing character recognition on the text image to be recognized to obtain a detection text having a plurality of characters in an arrangement order; The output unit is used to determine whether the detected text is a normal text or an inverted text according to the character encoding information of the characters in the preset character dictionary, and output the normal text of the detected text, and is also used to take each character in the detected text as a target character, search the character encoding information of the target character in the character encoding information; determine whether the target character is a normal character encoding or an inverted character encoding according to the character encoding information of the target character; and count a first quantity value of the normal character encoding and a second quantity value of the inverted character encoding in the detected text; The character encoding information of the characters in the character dictionary includes character encoding information of first-type characters, second-type characters, third-type characters and fourth-type characters predefined in a preset character set, the first-type characters are characters whose forward characters and inverted characters are different, the second-type characters are characters whose inverted characters are similar to forward characters of other characters in the character set or characters whose forward characters are similar to inverted characters of other characters in the character dictionary, the third-type characters are characters whose forward characters and inverted characters are the same, and the fourth-type characters are characters whose inverted characters are the same as forward characters of other characters in the character set; the first-type characters are characters whose forward characters and inverted characters are the same. The character encoding information of the first type of characters, the second type of characters, the third type of characters and the fourth type of characters includes pre-set forward character encoding, inverted character encoding and common character encoding information, the forward character encoding information is the character encoding information of the forward characters of the first type of characters and the second type of characters, the inverted character encoding information is the character encoding information of the inverted characters of the first type of characters and the second type of characters, the common character encoding information is the character encoding information of the third type of characters or the fourth type of characters, each character in the third type of characters and its corresponding character use the same character encoding information, and each character in the fourth type of characters and its corresponding character use the same character encoding information.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the text character detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image correction method, device and equipment and storage medium

    CN110647882A

  • Apparatus and method for character recognition and program thereof

    US20050053282A1