A character recognition method, device, electronic equipment and storage medium

By recognizing and acquiring character information and layout position parameters in the image to be recognized, the problem of layout information not being able to be restored in optical character recognition is solved, achieving efficient character recognition and layout preservation.

CN115546816BActive Publication Date: 2026-04-07ZHUHAI KINGSOFT OFFICE SOFTWARE +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing optical character recognition technology cannot effectively identify and restore the layout information of the characters to be recognized, resulting in text misalignment in the character recognition results.

Method used

By acquiring the image to be recognized, character information is identified and text character and layout position parameters are obtained. The receptive field is used to move and recognize characters in the image, generating characters in editable text format while maintaining layout consistency.

Benefits of technology

It achieves the simultaneous recognition of characters and acquisition of the layout position parameters of text characters, meeting users' editing and layout restoration needs, with high recognition efficiency and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546816B_ABST
    Figure CN115546816B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a character recognition method and device, electronic equipment and a storage medium. The method comprises: obtaining an image to be recognized; recognizing the image to be recognized to determine character information in the image to be recognized, obtaining text characters and layout position parameters of the text characters according to the character information; and setting the text characters in an editable text format according to the layout position parameters of the text characters. In the process of recognizing the image to be recognized, the edge positions of the text characters are determined by using intermediate data of a related algorithm, without additional edge detection of the text characters or an additional independent positioning algorithm. Thus, the recognized characters can meet the editing requirements of a user and the requirements of the user for layout restoration, and the recognition efficiency is high and the position determination cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to, but is not limited to, the technical field of character recognition, and in particular to a character recognition method and device, an electronic device, and a storage medium. BACKGROUND

[0002] Optical Character Recognition (OCR) technology (hereinafter referred to as "OCR technology") is a technology that converts printed characters to be recognized into a black-and-white dot matrix image file by using an optical method, and converts the characters in the image into a text format by recognition software, for further editing and processing by word processing software.

[0003] OCR technology can be used to recognize a single character to be recognized, or a batch of characters to be recognized, and the character recognition result obtained will be displayed in an editable text format. Although the above format converts the characters from an uneditable state to an editable state, it cannot restore the layout of the characters to be recognized, resulting in the problem of text running in the character recognition result. Therefore, a technology for recognizing layout information of characters to be recognized is needed. SUMMARY

[0004] To solve the above technical problems, the present application provides a character recognition method, device, electronic device and storage medium, which can recognize the layout position information of the characters to be recognized at the same time as recognizing the characters to be recognized, to generate the corresponding character recognition result.

[0005] Specifically, the present application is implemented by the following technical solutions:

[0006] In a first aspect, a character recognition method is provided, comprising:

[0007] obtaining an image to be recognized;

[0008] recognizing the image to be recognized to determine character information in the image to be recognized, and obtaining text characters and layout position parameters of the text characters according to the character information;

[0009] setting the text characters in an editable text format according to the layout position parameters of the text characters.

[0010] Optionally, the step of recognizing the image to be recognized to determine character information in the image to be recognized comprises:

[0011] setting a receptive field for the image to be recognized;

[0012] judging the character coverage range of the receptive field in the image to be recognized;

[0013] in a case where it is determined that the receptive field covers a complete character in the image to be recognized, recognizing a text character of the complete character covered by the receptive field, and a layout position parameter of the text character;

[0014] in a case where it is determined that the receptive field does not cover a complete character in the image to be recognized, moving the receptive field by a first moving step in a first direction.

[0015] Optionally, after the step of determining that the receptive field covers a complete character in the image to be recognized, the method further comprises:

[0016] determining whether the complete character covered by the receptive field contains an unrecognized character;

[0017] in a case where it is determined that the complete character covered by the receptive field contains an unrecognized character, recognizing a text character of the unrecognized character in the complete character covered by the receptive field, and a layout position parameter of the text character;

[0018] in a case where it is determined that the complete character covered by the receptive field does not contain an unrecognized character, moving the receptive field by a first moving step in a first direction.

[0019] Optionally, the step of, in a case where it is determined that the receptive field covers a complete character in the image to be recognized, recognizing a text character of the complete character covered by the receptive field, and a layout position parameter of the text character, comprises:

[0020] determining a pixel region covered by the complete character covered by the receptive field in the image to be recognized;

[0021] generating a corresponding text character according to pixels of the pixel region corresponding to the complete character;

[0022] determining a coordinate position of the text character, and taking the coordinate position as a layout position parameter of the text character.

[0023] Optionally, the step of determining the coordinate position of the text character comprises:

[0024] determining one or more slices corresponding to the text character in the receptive field, determining a slice corresponding to the text character in the receptive field as a central region of the text character, and respectively expanding the central region in a second direction and a third direction according to a first expansion step;

[0025] in a case where an edge of the central region expanded in the second direction satisfies a preset edge search termination condition, stopping the continuous expansion, and determining a position of the edge of the central region expanded in the second direction as a coordinate position of the text character in the second direction;

[0026] In a case where an edge of the center region expanded in the third direction satisfies a preset edge search termination condition, the expansion is stopped, and a position of the edge of the center region expanded in the third direction is determined as a coordinate position of the text character in the third direction.

[0027] Optionally, the preset edge search termination condition at least includes one of the following conditions:

[0028] Condition 1: a pixel disappears at an edge of the center region expanded in a corresponding expansion direction;

[0029] Condition 2: a receptive field corresponding to a slice to which the center region expanded in the corresponding expansion direction belongs is identified as covering a character that is a jump character or a split character;

[0030] Condition 3: a receptive field corresponding to a slice to which the center region expanded in the corresponding expansion direction belongs is identified as covering a character that is another character different from the text character.

[0031] Optionally, the identifying the to-be-identified image to determine character information in the to-be-identified image comprises:

[0032] extracting a character feature in the to-be-identified image, and converting the character feature into a character feature sequence;

[0033] identifying and decoding the character feature sequence to obtain a text character corresponding to the character feature sequence;

[0034] calculating a coordinate position of the text character, and taking the coordinate position as a layout position parameter of the text character.

[0035] Optionally, the to-be-identified image includes one character, or multiple characters, or a whole text; and in a case where the to-be-identified image includes the whole text, the whole text includes multiple adjacent characters.

[0036] Optionally, the identifying the to-be-identified image to determine character information in the to-be-identified image comprises:

[0037] identifying the to-be-identified image to determine character information of a whole text in the to-be-identified image; or identifying the to-be-identified image to determine character information of a preset number of characters in the to-be-identified image; or identifying the to-be-identified image to determine character information of characters of a preset language type in the to-be-identified image.

[0038] In a second aspect, a character recognition device is provided, comprising:

[0039] an obtaining module configured to obtain a to-be-identified image;

[0040] a recognition module configured to recognize the image to be recognized, determine character information in the image to be recognized, and acquire text characters and layout position parameters of the text characters according to the character information;

[0041] a conversion module configured to set the text characters in an editable text format according to the layout position parameters of the text characters.

[0042] In a third aspect, an electronic device is provided, which includes a memory and a processor, the memory storing a computer program for character recognition, and the processor is configured to read and run the computer program for character recognition to perform any of the above character recognition methods.

[0043] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is configured to perform any of the above character recognition methods when run.

[0044] The technical solutions provided by the embodiments of the present application can have the following beneficial effects:

[0045] In the process of recognizing the image to be recognized, the embodiments of the present application can not only recognize the text characters of the characters to be recognized in the image to be recognized, but also acquire the layout position parameters of the text characters, without additional edge detection of the characters to be recognized or additional independent positioning algorithms. Thus, the recognized characters to be recognized can meet the editing requirements of the user and the requirements of the user for layout restoration, and the recognition efficiency is high and the position determination cost is low.

[0046] Other features and advantages of the present application will be described in the following description, and some will become apparent from the description, or will be learned through practice of the present application. The objects and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0047] The accompanying drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification, and are used to explain the technical solutions of the present application together with the embodiments of the present application, and do not constitute a limitation on the technical solutions of the present application.

[0048] Figure 1 A character recognition method flowchart is shown for an exemplary embodiment of the present application;

[0049] Figure 2 Another character recognition method flowchart is shown for an exemplary embodiment of the present application;

[0050] Figure 3This is a schematic diagram of a slice and a sliding receptive field, illustrating an exemplary embodiment of the present invention.

[0051] Figure 4 This is a schematic diagram illustrating character edge search as an exemplary embodiment of the present invention;

[0052] Figure 5 A flowchart illustrating another character recognition method as an exemplary embodiment of the present invention;

[0053] Figure 6 This is a schematic diagram illustrating another character recognition effect as an exemplary embodiment of the present invention;

[0054] Figure 7 This is a schematic diagram illustrating the structure of a character recognition device according to an exemplary embodiment of the present invention. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

[0056] The steps illustrated in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented here.

[0057] First, the definitions of the various technical features involved in the embodiments of the present invention are listed below:

[0058] A slice refers to an image region unit with a preset pixel size. The size of the preset pixel can be adjusted and set according to different usage needs.

[0059] Receptive field: refers to the character recognition image area formed by connecting a preset number of slices one by one. For example, the corresponding character recognition image area is formed by connecting a preset number of slices end to end. The preset number of slices for the receptive field can be adjusted and set according to different usage needs.

[0060] An exemplary embodiment of the present invention illustrates a character recognition method, the process of which is as follows: Figure 1 As shown, the method includes the following steps:

[0061] Step 101: Obtain the image to be recognized;

[0062] Step 102: Identify the image to be identified, determine the character information in the image to be identified, and obtain the text character and the layout position parameters of the text character based on the character information;

[0063] Step 103: Set the text characters in the editable text format according to the layout position parameters of the text characters.

[0064] In one exemplary embodiment, in step 101, the image to be recognized contains a character to be recognized. The character to be recognized can originate from paper documents, such as various tickets, newspapers, books, manuscripts, and other printed materials; or from other forms of physical media, such as billboards, posters, road signs, and plaques; or from electronic files, such as image files in formats like PDF, JPG, and PNG. When the character to be recognized originates from paper documents or other forms of physical media, an image containing the character to be recognized can be acquired using a scanner, digital camera, or other device as the image to be recognized.

[0065] For example, a scanner can be used to scan the area corresponding to the character to be identified on a paper document or other physical medium, generating a scanned image containing the character to be identified, and this scanned image can be used as the image to be identified. Alternatively, a digital camera can be used to photograph the area corresponding to the character to be identified on a paper document or other physical medium, generating a photographic image (i.e., a photo) containing the character to be identified, and this photographic image can be used as the image to be identified. Alternatively, a photograph can be taken of the area corresponding to the character to be identified on a paper document or other physical medium, generating a video containing the character to be identified, and a screenshot of the interface containing the character in the video can be extracted and used as the image to be identified.

[0066] It should be noted that the characters to be recognized include one or more of the following types: Chinese characters, punctuation marks, English words, English letters, calculation symbols, and numbers.

[0067] In one exemplary embodiment, step 102 involves identifying characters in the image to be identified to determine character information. The character information includes characters in text format obtained through identification, and layout position parameters corresponding to the text-formatted characters. Text-formatted characters are simply text characters. The layout position parameters of the text characters include coordinates describing the start and end positions of the text characters in the image to be identified. In one exemplary embodiment, the coordinates are represented in pixels with the line direction as the axis; alternatively, other coordinate representation methods may be selected, and are not limited to a specific form.

[0068] In one exemplary embodiment, there are multiple ways to obtain character information of the character to be recognized. One way is to determine the character to be recognized in the image to be recognized, recognize only the character to be recognized, and obtain the character information of the character to be recognized, that is, to obtain the text format character of the character to be recognized and the layout position parameter corresponding to the text format character.

[0069] Another method involves recognizing all characters in the image to be recognized, obtaining character information for all characters in the image, and then selecting the character information corresponding to the character to be recognized. This yields the text-formatted character of the character to be recognized, along with the corresponding layout position parameters. The character to be recognized can be manually set by the user, for example, by using the terminal to select a portion of the characters in the image as the character to be recognized; alternatively, the character to be recognized can be set by default, for example, by defaulting to all characters in the image. This solution does not impose any restrictions on this.

[0070] In another exemplary embodiment, there are several ways to identify characters in an image to be identified. One method is to identify the text character of each character in the image to be identified and the layout position parameter of that character. Another method is to identify all the text characters of the character to be identified in the image, as well as the start layout position parameter corresponding to the first character of the character to be identified and the end layout position parameter corresponding to the last character of the character to be identified. Yet another method is to identify the text characters of a preset number of characters in the image to be identified and the corresponding layout position parameters of that preset number of characters. This solution does not limit these methods.

[0071] In one exemplary embodiment, in step 103, the text character is set in an editable text format according to the layout position parameters of the text character. That is, based on the text character and its layout position parameters, the format of the character to be recognized is converted into an editable text format consistent with the layout of the image to be recognized. With the text character and its layout position parameters obtained, the format of the character to be recognized is converted based on the text character and its layout position parameters. In this way, while converting the format of the character to be recognized into an editable format, the layout of the character to be recognized is made consistent with the layout of the character to be recognized in the image to be recognized. Here, "layout consistency" means that the characters and positions of the newly generated text format are the same as the corresponding characters and positions in the image to be recognized. In this way, the recognized characters can satisfy both the user's editing needs and the user's need to maintain consistent layout (i.e., layout restoration). Furthermore, the acquisition of the text character to be recognized and the acquisition of the layout position parameters corresponding to the text character are completed in the same step, without the need for additional positioning operations. That is, there is no need for additional operations related to edge detection position of the character to be recognized or additional independent positioning algorithms. This makes the recognition efficiency high and the position determination cost low.

[0072] It should be noted that the starting position of the characters to be recognized in the editable text format can be any position in the editable document, such as the first line, the fifth line, etc. of the editable document, or, for example, the user touches the display screen with a finger or clicks on the corresponding position of the editable document with a mouse. When determining the starting position of the characters to be recognized, when the text characters of a line cannot be correspondingly arranged in a line of the editable document, for example, the display direction of the editable document can be changed, such as changing the vertical display to a horizontal display, to increase the line length of the typeset text characters; in addition, the text characters can also be adjusted to be arranged at the beginning position of the next line of the line in the editable document where the arrangement starting position is located, and the beginning position refers to the position where text characters can be edited and input at the beginning of a line of the editable document; in addition, the character spacing between the text characters can be uniformly reduced by a preset ratio, for example, the character spacing between the text characters is reduced to one-third, one-fifth, one-half, etc. of the current character spacing, and the specific embodiments of the present invention do not limit this character spacing reduction ratio.

[0073] For example, when recognizing the characters to be recognized in an invoice bill, if both the text characters of the characters to be recognized and the layout position parameters corresponding to the text characters of the characters to be recognized are recognized, it will avoid the running of the recognized text characters, which is beneficial to, for example, the filling and editing of financial data, and improves the user's convenience experience.

[0074] Further, recognize the image to be recognized and determine the character information in the image to be recognized, including: setting a receptive field for the image to be recognized, that is, setting a receptive field for the characters to be recognized in the image to be recognized. According to the coverage of the receptive field on the characters to be recognized in the image to be recognized, an identification operation is performed. Among them, the receptive field covers the area corresponding to the characters to be recognized. The receptive field can be formed by connecting any natural number of slices one by one, such as 6 slices, 5 slices, 12 slices, etc. The number of slices of the receptive field can be set manually by the user or can be a default value, and the present invention does not limit this. Exemplarily, in the present invention, the receptive field is formed by connecting 8 slices one by one.

[0075] Judge the character coverage range of the receptive field in the image to be recognized. When it is determined that the receptive field covers the complete character in the image to be recognized, use the receptive field to recognize the above complete character, and obtain the text character corresponding to the above complete character according to the recognition result. For example, if the above complete character is an uneditable image format of "多", after recognition, the editable text format of "多" can be obtained. In addition, the receptive field can also be used to obtain the layout position parameters of the text character corresponding to the above complete character. When it is determined that the receptive field does not cover the complete character in the image to be recognized, move the receptive field along the first direction by the first moving step.

[0076] The "first direction" refers to the direction of the character line to be recognized, that is, the direction in which the receptive field moves along the text line of the character to be recognized. Alternatively, it can be based on the direction in which the character is displayed normally and accurately by the naked eye when the user is looking at it at eye level. The aforementioned first direction can be understood as a horizontal direction parallel to the ground. In addition, the first movement step size can be the number of pixels corresponding to a preset number of slices. In a preferred embodiment of the present invention, the first movement step size refers to the number of pixels corresponding to one slice, that is, the number of pixels corresponding to the width of one slice that the receptive field moves each time.

[0077] It should be noted that after the receptive field moves by the first step along the first direction, the step of determining the character coverage area of ​​the receptive field in the image to be recognized continues. The specific method is described above and will not be repeated here.

[0078] If the receptive field determines that a complete character is covered in the image to be recognized, it is necessary to further determine whether the complete character covered by the receptive field contains unrecognized characters, that is, whether the complete character currently covered by the receptive field has already undergone character recognition processing. If it is determined that the complete character covered by the receptive field contains unrecognized characters, the unrecognized characters within the complete character covered by the receptive field are recognized. After recognition, the text character of the unrecognized character and its layout position parameters are obtained. The complete character covered by the receptive field may contain some unrecognized characters and some recognized characters. In this case, only the unrecognized characters are processed for character recognition. The progress of the recognition processing of the complete character covered by the receptive field can be determined based on the text character information corresponding to the already recognized complete characters.

[0079] For example, if the receptive field covers five complete characters in the image to be recognized, and two of the complete characters have been identified as text characters, then the remaining three complete characters are determined to be unrecognized characters. Character recognition processing is only required for these three complete characters. If no text characters correspond to the identified complete characters, then all complete characters covered by the receptive field are determined to have not undergone character recognition processing. Character recognition processing must be performed on all of these complete characters. Conversely, if it is determined that the complete characters covered by the receptive field do not contain unrecognized characters, the receptive field is moved along the first direction by a first moving step. The specific moving direction and moving pixels of the receptive field have been explained above and will not be repeated here.

[0080] When it is determined that the receptive field covers a complete character in the image to be recognized, the text character of the complete character covered by the receptive field is identified using the layout position parameters of the text character. This includes: determining the pixel area covered by the complete character in the image to be recognized. Specifically, the determination of the complete character covered by the receptive field in the image to be recognized is achieved as follows: identifying the pixel area covered by the receptive field in the image to be recognized along a first direction, and obtaining the pixel range of the pixel area covered by the receptive field. That is, determining the range of the pixel-covered area in the current receptive field. The pixel-covered area corresponds to the strokes of the character.

[0081] When it is determined that the receptive field covers a complete character in the image to be recognized, the method for recognizing the text character of the complete character covered by the receptive field, using the layout position parameter of the text character, further includes: generating the corresponding text character based on the pixel composition of the pixel region corresponding to the complete character, and determining the coordinate position of the text character based on the pixel composition of the aforementioned pixel region, using this coordinate position as the layout position parameter of the text character. Here, generating the corresponding text character refers to setting the corresponding text character based on the pixels of the pixel region corresponding to the complete character.

[0082] It should be noted that after all characters in the text line containing the character to be recognized have been recognized, the receptive field will then be deployed to the next adjacent text line containing the character to be recognized, and the aforementioned steps for recognizing the character to be recognized will be performed on that text line.

[0083] Further, determining the coordinate positions of the text characters includes:

[0084] Determine one or more slices in the receptive field corresponding to the text character, and define the slice in the receptive field corresponding to the text character as the central region of the text character. Expand the central region in the second direction and the third direction respectively according to the first expansion step.

[0085] Determine whether the edge of the central region expanding along the second direction meets the preset edge search termination condition; if the preset edge search termination condition is met, stop expanding further, and determine the edge position of the currently expanded central region in the second direction as the coordinate position of the text character in the second direction. In other words, if the edge of the central region expanding along the second direction meets the preset edge search termination condition, stop expanding further, and determine the edge position of the currently expanded central region in the second direction as the coordinate position of the text character in the second direction.

[0086] Furthermore, it determines whether the edge of the central region expanding along a third direction satisfies a preset edge search termination condition; if the preset edge search termination condition is met, it stops expanding further and determines the edge position of the currently expanded central region in the third direction as the coordinate position of the text character in the third direction. In other words, if the edge of the central region expanding along a third direction satisfies the preset edge search termination condition, it stops expanding further and determines the edge position of the currently expanded central region in the third direction as the coordinate position of the text character in the third direction.

[0087] The "second direction" refers to the direction along the character line, from the center area of ​​the text character towards the beginning of the text character; the "third direction" refers to the direction along the character line, from the center area of ​​the text character towards the end of the text character. The first expansion step size refers to the pixel width of each expansion of the center area, that is, the center area expands its edges along the second and third directions according to the first expansion step size. After expansion, it is determined whether the preset edge search termination condition has been met. If it has, the expansion in that direction is stopped, indicating that the edge position of the text character in that expansion direction has been determined, and the coordinates of this edge position are used as the coordinate position of the text character in that expansion direction.

[0088] After performing the edge search of the text character along the second and third directions as described above, the edge position of the text character in these two directions can be determined, and the coordinates of the edge position of the text character in these two directions can be used as the coordinate position of the text character.

[0089] In one exemplary embodiment, the above steps determine the coordinate positions of the left and right edges of the text character, that is, determine the starting and ending edge positions of the text character. Here, "left edge" and "right edge" can be understood as the two edges along the horizontal direction of the character, based on the direction in which the character is normally and accurately displayed by the naked eye when viewed at eye level.

[0090] In one exemplary embodiment, step 102 includes:

[0091] A convolutional neural network is used to extract character features from the image to be recognized, and the character features are converted into a character feature sequence.

[0092] Identify and decode character feature sequences to determine each text character included in the image to be identified.

[0093] Calculate the coordinate position of a text character and use it as a layout position parameter for the text character. In one exemplary embodiment, calculating the coordinate position of a text character includes: using the receptive field corresponding to the text character to determine the coordinate position of the text character in each preset direction. Specifically, taking a slice of the preset position in the receptive field corresponding to the text character as the central region, the central region is sequentially expanded in the preset direction according to a first expansion step size until it is determined that the edge of the expanded central region in the expansion direction satisfies a preset edge search termination condition, at which point the expansion stops; based on the edge position of the currently expanded central region in the expansion direction, the edge position of the character in the preset direction is determined.

[0094] In one exemplary embodiment, the preset direction may include one direction, which means determining the edge position of the text character in one direction. For example, if the preset direction is a second direction (to the left), then the coordinate position of the left edge of the text character is determined as the layout position parameter of the text character. Alternatively, the preset direction may include two directions, which means determining the edge position of the text character in two directions. For example, if the preset direction is a second direction and a third direction (to the left and right), then the coordinate positions of the left and right edges of the text character are determined and used together as the layout position parameter of the text character, that is, the layout position parameter of the text character includes the left starting coordinate position and the right ending coordinate position of the text character.

[0095] In one exemplary embodiment, the preset direction is the line direction of the characters in the image to be recognized, generally the reference direction for character recognition. For example, if the left-right direction is used as the reference direction for character recognition, then one preset direction can include: a leftward direction or a rightward direction; two preset directions can include: a leftward direction and a rightward direction. If the reference direction for character recognition is the up-down direction, then one preset direction can include: an upward direction or a downward direction; two preset directions can include: an upward direction and a downward direction. Here, "upward direction" and "downward direction" can be understood as the direction along the vertical axis of the character, based on the direction in which the character is normally and accurately displayed by the naked eye when the user is looking at it at eye level.

[0096] For example, if the coordinates of the beginning of an editable document correspond to the left edge of the image to be recognized, then the coordinates of each edge of the text character can be calculated starting from the beginning of the document. The coordinates of the text character can be represented in pixels. Alternatively, a predetermined position in the image to be recognized can be used as the starting point for the coordinate calculation.

[0097] According to the embodiment of the present invention, character recognition is performed. When text characters are recognized, their position information is determined. Based on this recognition result, both the text characters and their corresponding layout information are obtained. This ensures that the recognized characters meet both the user's editing needs and their need for layout restoration, resulting in high recognition efficiency and low position determination cost.

[0098] In one exemplary embodiment, when converting the format of the character to be recognized into an editable text format consistent with the layout of the image to be recognized in step 103, the editable text format can be an XML (Extensible Markup Language) file format, or a JSON (JavaScript Object Notation) format; or other correspondingly defined rich text structures. This is not limited to the content exemplified in this invention.

[0099] Furthermore, the preset edge search termination condition includes at least one of the following conditions:

[0100] Condition 1: The pixels in the expanded central region disappear at the edge of the corresponding expansion direction.

[0101] Condition 2: The receptive field of the slice corresponding to the edge of the expanded central region in the corresponding expansion direction is identified as covering characters that are either transition characters or segmentation characters.

[0102] Condition 3: The receptive field of the slice corresponding to the edge of the expanded central region in the corresponding expansion direction is identified as covering a character that is different from the text characters mentioned above.

[0103] Conditions 1, 2, and 3 above are used to distinguish different edge search termination conditions and do not represent the priority or execution order of the conditions.

[0104] In one exemplary embodiment, when expansion stops due to condition 1, determining the edge position of the character in the expansion direction based on the edge position of the currently expanded central region in that expansion direction includes:

[0105] The edge position of the currently expanded central region in the expansion direction is determined as the edge position of the text character in the expansion direction; that is, the edge position of the currently expanded central region in the expansion direction is taken as the edge position of the text character in the expansion direction.

[0106] In one exemplary embodiment, when expansion stops due to condition 2, determining the edge position of the character in the expansion direction based on the edge position of the currently expanded central region in that expansion direction includes:

[0107] Determine the slice to which the currently expanded central region belongs at the edge of the expansion direction, and use the edge position of the slice to which it belongs at the edge of the expansion direction as the edge position of the character at the edge of the expansion direction.

[0108] In one exemplary embodiment, when expansion stops upon satisfying condition 3, determining the edge position of the character in the expansion direction based on the edge position of the currently expanded central region in that expansion direction includes:

[0109] Determine the slice to which the currently expanded central region belongs at the edge of the expansion direction, and use the edge position of the slice in the opposite direction of the expansion direction as the edge position of the character in that expansion direction. For example, if the expansion direction is to the left, then the right edge position of the slice is the left edge position of the text character; and so on for other expansion directions.

[0110] In one exemplary embodiment, the preset direction includes a second direction and a third direction; for example, left and right. For a text character, the coordinate positions of its left and right edges are determined. The edge to the left of the central region is also called the left edge of the central region, and the edge to the right of the central region is also called the right edge of the central region; the edge to the left of a text character is also called the left edge of the character, and the edge to the right of a character is also called the right edge of the character.

[0111] Determining the left-hand edge position of a text character based on its receptive field includes:

[0112] Using the slice at a preset position in the receptive field corresponding to the text character as the central region, the central region is expanded to the left sequentially according to the first expansion step size until the expanded central region satisfies preset condition 2 at its left edge, at which point the expansion stops. The slice to which the left edge of the currently expanded central region belongs is determined, and the left edge position of the text character is taken as the left edge position of the slice to which it belongs. That is, the slice to which the left edge of the currently expanded central region belongs is determined, and the left edge position of the slice to which it belongs is taken as the left edge position of the text character. Alternatively, using the slice at the preset position in the receptive field corresponding to the text character as the central region, the central region is expanded to the left sequentially according to the first expansion step size until the expanded central region satisfies preset condition 3 at its left edge, at which point the expansion stops. The slice to which the left edge of the currently expanded central region belongs is determined, and the right edge position of the slice to which it belongs is taken as the left edge position of the text character. That is, the slice to which the left edge of the currently expanded central region belongs is determined, and the right edge position of the slice to which it belongs is taken as the left edge position of the text character.

[0113] Based on a text character and its corresponding receptive field, the right edge position of the text character can be determined as follows:

[0114] Using the slice at a preset position within the receptive field corresponding to the text character as the central region, the central region is expanded sequentially to the right according to the first expansion step size until the expanded central region satisfies preset condition 2 at its right edge, at which point the expansion stops. The slice to which the right edge of the currently expanded central region belongs is determined, and the right edge position of that slice is taken as the left edge position of the character. That is, the slice to which the right edge of the currently expanded central region belongs is determined, and the right edge position of that slice is taken as the right edge position of the text character. Alternatively, using the slice at a preset position within the receptive field corresponding to the text character as the central region, the central region is expanded sequentially to the right according to the first expansion step size until the expanded central region satisfies preset condition 3 at its right edge, at which point the expansion stops. The slice to which the right edge of the currently expanded central region belongs is determined, and the left edge position of that slice is taken as the right edge position of the character. That is, the slice to which the left edge of the currently expanded central region belongs is determined, and the left edge position of that slice is taken as the right edge position of the character.

[0115] In one exemplary embodiment, the slice at a predetermined location in the receptive field includes:

[0116] One or more adjacent slices in the middle of the receptive field.

[0117] For example, the receptive field of a text character A consists of 8 slices, numbered 1-8. The initial central region can be the 3rd-4th slice, the 3rd-5th slice, the 4th-5th slice, or the 3rd slice.

[0118] Further, the process of identifying the image to be identified and determining the character information in the image to be identified includes: extracting character features from the image to be identified and converting the character features into a character feature sequence.

[0119] Determining character information in an image to be recognized also includes: recognizing and decoding character feature sequences to obtain the text characters corresponding to the character feature sequences.

[0120] Determining character information in the image to be recognized also includes: calculating the coordinate positions of the text characters and using the coordinate positions as layout position parameters of the text characters.

[0121] Furthermore, the image to be recognized includes: one character, multiple characters, or an entire text segment. Specifically, the character to be recognized in the image includes: one character, multiple characters, or an entire text segment; wherein, when the image to be recognized includes an entire text segment, the entire text segment includes multiple adjacent characters.

[0122] Further, the image to be recognized is identified, and character information in the image to be recognized is determined, including:

[0123] The method involves identifying an image to be recognized and determining the character information of the entire text segment within that image; or, identifying an image to be recognized and determining the character information of a preset number of characters within that image; or, identifying an image to be recognized and determining the character information of characters of a preset language type within that image. It should be noted that, depending on the application requirements, the character recognition method provided in this disclosure can obtain the positional parameters of individual text characters and layout in the image to be recognized, as well as an editable text format that matches the layout of the image to be recognized after conversion.

[0124] Alternatively, a preset number of text characters and layout position parameters can be obtained, as well as an editable text format that is consistent with the layout of the image to be recognized after conversion; wherein, the layout position parameters corresponding to the multiple text characters can be the layout position parameters of each text character individually, or the layout position parameters of adjacent text characters after merging.

[0125] This invention also provides a character recognition method, such as... Figure 2 As shown, the method includes:

[0126] Steps 201-204, wherein steps 201-203 are consistent with steps 101-103 in the aforementioned embodiment.

[0127] Step 204: Determine at least one text segment and the layout position parameter corresponding to the text segment based on the editable text format that is consistent with the layout of the image to be recognized.

[0128] The text segment consists of multiple adjacent characters.

[0129] In one exemplary embodiment, step 204 includes:

[0130] Semantic analysis is performed on some or all of the identified text characters to split them into at least one text segment; the layout position parameters of each text segment are determined based on the text characters contained in each text segment and the layout position parameters corresponding to the text characters.

[0131] This exemplary embodiment can perform semantic segmentation on some or all of the identified text strings according to semantics, forming text segments such as words, phrases, short sentences, etc., and determine the layout position parameters of each text segment based on the text characters included in each text segment and the layout position parameters of each text character.

[0132] In one exemplary embodiment, step 204 includes:

[0133] Based on the language type of each text character in the identified text characters, adjacent text characters in part or all of the identified text characters are grouped into at least one text segment according to the language type grouping rules; based on the text characters contained in each text segment and the layout position parameters corresponding to the characters, the layout position parameters of each text segment are determined; wherein, the language types include: Chinese, English, numbers and punctuation.

[0134] This exemplary embodiment can group the identified partial or complete text strings according to language type, form text segments such as Chinese, English, numbers, numbers + punctuation, and / or Chinese + punctuation by adjacent text characters, and determine the layout position parameters of each text segment based on the text characters included in each text segment and the layout position parameters of each text character.

[0135] In one exemplary embodiment, after performing step 204, the method further includes: step 205, saving at least one text segment and the layout position parameters of the corresponding text segment.

[0136] In one exemplary embodiment, an editable text format consistent with the layout of the image to be recognized, as used in step 103, is employed to save at least one text segment and the layout position parameters of the corresponding text segment.

[0137] In one exemplary embodiment, the first expansion step size is in pixels. For example, the first expansion step size is 1 pixel. Taking a preset direction or a second direction as left as an example, the coordinate position of the text character to the left is determined based on the receptive field corresponding to the text character.

[0138] Using the slice at the preset position in the receptive field corresponding to the text character as the central region, the expansion is carried out to the left by one pixel each time to form the expanded central region. The expansion stops when the left edge of the expanded central region meets the preset edge search termination condition. Based on the current left edge position of the expanded central region, the coordinate position of the character on the left is determined.

[0139] In one exemplary embodiment, the characters include: Chinese characters, words, letters, numbers, punctuation marks, etc.

[0140] In one exemplary embodiment, step 102 or 202 includes:

[0141] Based on the image to be recognized, a scheme of convolutional recurrent neural network + connection time classifier is used to identify the character information in the image.

[0142] A convolutional recurrent neural network (CTC) combined with a connection-temporal classifier (CT) is employed for character recognition. First, features are extracted using a convolutional network. Then, contextual features are extracted using a recurrent network layer, followed by a sequence recognition layer to identify possible categories and their probabilities. Next, the sequence is decoded by a connection-temporal classifier (CTC, also known as a Greedy Decoder). This decoder merges duplicate characters to obtain the actual string to be recognized. This scheme can recognize strings in the image and obtain the receptive field for each character. The receptive field for each character is intermediate data in the execution process of the CTC scheme and is considered an algorithmic attribute of the scheme. Those skilled in the art will understand this aspect, but it is not within the scope of this application. Further details are omitted here.

[0143] In one exemplary embodiment, step 102 or 202 includes:

[0144] Based on the image to be recognized, a scheme of convolutional neural network + connection time classifier is used to identify the character information in the image.

[0145] To improve execution efficiency and reduce the overall resource requirements of the above embodiments, the convolutional recurrent neural network (CTC) + connection-temporal classifier (CTG) scheme is improved by removing the recurrent network layer and adopting a CTC + CTG scheme for character recognition: First, features are extracted through the convolutional network, then the sequence recognition layer identifies the possible categories and their probabilities; then, the sequence is decoded by the connection-temporal classifier (CTC, also known as a Greedy Decoder) to obtain the actual string to be recognized. This scheme can recognize the string in the image to be recognized and obtain the receptive field corresponding to each character.

[0146] The character recognition scheme provided in this embodiment of the invention utilizes attributes or intermediate data in relevant character recognition algorithms to determine the position information of each character during the string recognition process.

[0147] Example 1

[0148] The solution disclosed in this invention performs character localization during the character recognition process. It determines the layout position parameters of each text character, also known as the character's position information, to achieve character localization.

[0149] Before describing the process in this example, let's clarify a few concepts:

[0150] Slice: Refers to a virtual unit for dividing the width of an image. In this example, a slice is 8 pixels wide and 32 pixels high. Figure 3and Figure 4 As shown in the example. The size of the slice can be adjusted according to the image to be identified, and is not limited to this example.

[0151] Receptive field: In this example, it refers to the character recognition image region (width, 8*8 = 64 pixels) formed by connecting 8 slices sequentially. In convolutional recurrent neural networks or convolutional neural network schemes, it is also the size of the actual input to the convolutional neural network for feature extraction. The size of the receptive field can be adjusted according to the image to be recognized, and is not limited to this example.

[0152] Sliding receptive field: A fundamental concept in deep learning, representing the pixel region currently being computed. It slides according to the size of the sliding window, with each sliding position corresponding to a sliding receptive field.

[0153] Rich text: A custom structure for expressing OCR results, which includes not only the recognized characters but also location information; alternatively, it may contain extended additional text information, such as Chinese-English word separation.

[0154] This example provides a character recognition method, as follows:

[0155] The preset directions include two: left and right, that is, each character corresponds to the determination of the left edge position and the right edge position; the initial position of the central region is the 4th-5th slice of each receptive field; the first expansion step size is 1 pixel, that is, expanding left or right pixel by pixel to perform edge search.

[0156] The methods for identifying this character location include:

[0157] Step a: Obtain the image to be recognized.

[0158] Step b: Identify the image to be identified, determine the character information in the image to be identified, and obtain the text characters and their layout position parameters based on the character information.

[0159] Step c: Set the text characters in the editable text format according to the layout position parameters of the text characters.

[0160] Step ac is implemented with reference to steps 101-103 and related aspects of the aforementioned embodiments.

[0161] like Figure 3 As shown, the image to be recognized contains a character to be recognized. The character to be recognized in the image is "multi-point breakthrough", which means active support.

[0162] Step b includes: setting a receptive field for the image to be recognized for character recognition. When the receptive field is set to 1 - 8 slices, the text characters "多" and the layout position parameters corresponding to "多" are recognized; when the receptive field is set to 5 - 12 slices, the text characters "点" and the layout position parameters corresponding to "点" are recognized.

[0163] Among them, in step b, the coordinate position of the text character is determined in the string, such as Figure 5 shown, including:

[0164] Step 5021: According to the receptive field corresponding to each character, determine the coordinate position of the left edge of each character.

[0165] Step 5022: According to the receptive field corresponding to each character, determine the coordinate position of the right edge of each character.

[0166] The coordinate position of the left edge and the coordinate position of the right edge of each character constitute the layout position parameter of each text character.

[0167] As Figure 4 shown, taking the character "多" as an example, the fourth and fifth slices from the left cover many strokes of the character "多", enabling the classification network to recognize that these two slices are the top 2 slices with the highest probability of recognizing the character "多" in the receptive field of the character "多", that is, the probability of recognizing the character "多" based on these two slices is the greatest. It is deduced that when the input height is 32 pixels, the width of the sliding receptive field of the convolutional recurrent neural network or convolutional neural network is approximately 8 * 8 = 64 pixels, and the width of each slide is 8 pixels. Generally, for each slide position, the character with the highest probability of the output result of the corresponding receptive field is the character at the center of the receptive field. For example, Figure 4 for the character "多" in, it corresponds to the horizontal range of about 24 - 40 pixels from the left side. Using this rough range, through pixel edge search, more accurate starting and ending coordinates can be obtained.

[0168] For the character "多", its step 5021 includes:

[0169] The receptive field corresponding to the character "多" set in step b is 1 - 8 slices, with the 4th - 5th slices as the central area, that is, the initial central area. Expand 1 pixel width to the left each time, and determine whether the left edge of the expanded central area meets the preset edge search termination condition; that is, as Figure 4 shown, starting from the left edge of the 4th slice, the left edge (indicated by the left vertical line) moves one pixel to the left each time to probe where the left edge of the character "多" is.

[0170] When the preset edge search termination condition is satisfied, the left edge of the character "多" is determined according to the left edge of the currently expanded central region.

[0171] Among them, the preset edge search termination condition includes at least one of the following:

[0172] Condition 1: Pixels disappear at the edge of the currently expanded central region in the corresponding expansion direction.

[0173] Condition 2: The receptive field to which the slice to which the edge of the currently expanded central region belongs in the corresponding expansion direction belongs is recognized as covering a character that is a jump character or a segmentation character.

[0174] Condition 3: The receptive field to which the slice to which the edge of the currently expanded central region belongs in the corresponding expansion direction belongs is recognized as covering a character that is another character.

[0175] In the case where the image to be recognized is relatively clean and there is a large difference between the pixels of the character to be recognized and the background, if it is determined that pixels disappear at the edge of the expanded central region, it is determined that Condition 1 is satisfied, and the continued expansion (continued exploration) is stopped, and the edge of the currently expanded central region is used as the edge position of this character. For example, for the character "多", if Figure 4 pixels disappear at the edge shown by the left vertical line, the left edge of the character "多" is determined to be the coordinate position corresponding to the left vertical line.

[0176] In some implementations, when the characters to be recognized in the image to be recognized have strokes connected together, or are interfered by noise and stains, making it difficult to find the edges of the characters to be recognized, the boundary of the left jump character or segmentation character is used as the termination condition for edge search.

[0177] For example, assume that the receptive field corresponding to the 4th slice recognizes the character "多", while the receptive field corresponding to the 3rd slice recognizes a segmentation character. Then, if pixel search is performed from the left edge of the 4th slice to the left and no edge is found until the left end of the 3rd slice, the left edge of the character "多" will terminate at the left edge of the 3rd slice.

[0178] Or, assume that the receptive field corresponding to the 4th slice recognizes the character "多", while the recognition result of the 2nd slice is a left quotation mark, which is different from the character "多", that is, a jump occurs. Then, if pixel search is performed from the left edge of the 4th slice to the left and no edge is found until the right end of the 2nd slice, the left edge of the character "多" will terminate at the right edge of the 2nd slice.

[0179] For the character "多", its step 5022 includes:

[0180] The receptive field corresponding to the character "多" set in step b is 1 - 8 slices, with the 4th - 5th slices as the central area, that is, the initial central area. The width is expanded by 1 pixel to the right each time, and it is judged whether the right edge of the expanded central area meets the preset edge search termination condition; that is, as Figure 4 shown, starting from the right edge of the 5th slice, the right edge (indicated by the right vertical line) moves one pixel to the right each time to explore where the right edge of the character "多" is.

[0181] When the preset edge search termination condition is met, the right edge of the character "多" is determined according to the right edge of the currently expanded central area, which is similar to expanding to the left and judging whether the edge search termination condition is met.

[0182] In the case where the image to be recognized is relatively clean and there is a large difference between the pixels of the character to be recognized and the background, if it is determined that the pixels disappear at the edge of the expanded central area, that is, the right edge of the central area expands to Figure 4 the position shown by the right vertical line, it is determined that condition 1 is met, and the expansion continues to be stopped, and the edge of the currently expanded central area is used as the edge position of this character. Exemplarily, Figure 4 the position shown by the right vertical line and the right edge of the currently expanded central area are the right edge positions of the character "多".

[0183] In the case where the characters to be recognized in the image to be recognized are connected by strokes, or there are noises and stains interfering, continue to search to the right. The right edge of the central area expands to the left edge of the 7th slice. If it is judged that the recognition result of the receptive field to which the 7th slice belongs is the character "点", it is determined that condition 3 is met, and the left edge position of the 7th slice is used as the right edge position of the character "多".

[0184] In this example, the edge position is represented by coordinate information, indicating the distance of the determined edge from the leftmost (left edge) of the original picture. For example, as Figure 4 shown in the figure, the coordinate of the left edge of the character "多" is approximately 20, and the coordinate of the right edge is approximately 42, indicating that the left edge of the character "多" is about 20 pixels away from the leftmost of the picture, and the right edge of the character "多" is about 42 pixels away from the left side of the picture. Then, the position information constituting the character "多" is (20, 42), respectively representing the relative positions of the left and right edges of the character "多" compared to the starting point of the left edge of the image to be recognized.

[0185] In the case where there is interference in the image to be recognized, if the left and right edges of the character "多" meet condition 2 or condition 3, the search continues to be stopped. In this way, it is determined that the coordinate of the left edge of the character "多" is approximately 16, and the coordinate of the right edge is approximately 48. The positioning information constituting the character "多" is (16, 48), respectively representing the positions of the left and right edges of the character "多".

[0186] For other characters, refer to the processing steps of the above "duo" character, and respectively execute steps 5021-5023 to determine the coordinate positions corresponding to each text character one by one. No further examples will be given here.

[0187] This example solution utilizes the algorithmic attributes of the classification network solution. Starting from one side of the slice (central region, as shown in Figure 6 shown) in the receptive field where the text character is located, it tries to obtain the edge of the text pixel (accurate position) through trial and error, and uses the jumping character, space or segmentation character as one of the termination conditions for edge trial and error. Determine the positioning information of the text character one by one. During the character recognition process, it can utilize the intermediate attribute of the character recognition scheme related to the convolutional neural network + connectionist temporal classification or convolutional recurrent neural network + connectionist temporal classification - the receptive field corresponding to the character, for character positioning, so that the recognized character can not only meet the user's editing needs, but also meet the user's needs for layout restoration, with high recognition efficiency and low position determination cost. It improves the execution efficiency of the overall character recognition + character positioning solution, reduces the demand for computing resources, and reduces the device resource limitations of the overall solution application.

[0188] Example Two

[0189] This example provides a character recognition method. After determining the layout position parameters of the character, this method further includes:

[0190] According to the recognized text characters and the layout position parameters of the text characters, determine at least one text segment and the layout position parameters corresponding to the text segment; wherein, the text segment includes multiple adjacent characters.

[0191] In this example, determining at least one text segment and the layout position parameters corresponding to the text segment according to the recognized text characters and the layout position parameters of the text characters includes:

[0192] According to the language type of each recognized text character, group adjacent characters into at least one text segment according to the language type grouping rules; according to the text characters included in each text segment and the layout position parameters corresponding to the characters, determine the layout position parameters of each text segment; wherein, the language types include: Chinese, English, numbers, and punctuation.

[0193] For example, Figure 6 as shown, after respectively determining the layout position parameters for each text character, then according to the language attributes of each text character, split the recognized text string to determine 5 text segments, namely: Chinese segment, English segment, Chinese + punctuation segment, English segment, and Chinese + punctuation segment.

[0194] Based on the layout position parameters of the text characters contained in each text segment, determine the layout position parameters of this text segment. For example, for text segment 1 "Load and display the image at", use the left edge coordinate position of the character "把" as the left edge coordinate position (coordinate a) of text segment 1, and use the right edge coordinate position of the character "在" as the right edge coordinate position (coordinate d) of text segment 1, to form the layout position parameters (coordinate a, coordinate d) of text segment 1, representing the start and end positions of text segment 1. The example information of other text segments is determined similarly and will not be elaborated here.

[0195] In an exemplary embodiment, according to the positioning information of the above 5 text segments, perform splitting or coloring processing on these 5 text segments, and the result is as Figure 7 shown.

[0196] Example 3

[0197] In this example, a method for character recognition provides an editable text format that is consistent with the layout of the image to be recognized and saves it in a rich text format.

[0198] In an exemplary embodiment, it can be saved in a json or xml file. For example, an editable text format described in json is used. Specifically, set the recognition result of the overall text string to be saved; split it into multiple text segments (multiple strings) according to the differences between Chinese and English languages, and record the content and positioning information of each, corresponding to the start and end coordinates of multiple text strings on the x-axis of the original image; or record according to the splitting results and corresponding position information of characters, words, numbers, or punctuation marks.

[0199] Those skilled in the art can use other replacement methods or formats to save the text characters and the layout position parameters of the text characters recognized in the embodiments of the present invention. It is not limited to the method exemplified in the embodiments of the present invention.

[0200] The embodiments of the present invention also provide a character recognition device, as ​ shown, including:

[0201] An acquisition module 701, configured to acquire an image to be recognized;

[0202] A recognition module 702, configured to recognize the image to be recognized, determine the character information in the image to be recognized, and acquire the text characters and the layout position parameters of the text characters according to the character information;

[0203] A conversion module 703, configured to set the text characters in an editable text format according to the layout position parameters of the text characters.

[0204] This invention also provides an electronic device, including a memory and a processor. The memory stores a computer program for character recognition, and the processor is configured to read and run the computer program for character recognition to execute the character recognition method of any of the above embodiments.

[0205] This invention also provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the character recognition method of any of the above embodiments at runtime.

[0206] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

Claims

1. A character recognition method, characterized in that, include: Acquire the image to be recognized; A neural network is used to identify the image to be identified based on a set receptive field, determine the character information in the image to be identified, and obtain the text character and its layout position parameters based on the character information. Obtaining the layout position parameters of the text character includes: during the text character recognition process, determining that the text character corresponds to one or more slices in the receptive field, defining the slice corresponding to the text character in the receptive field as the central region of the text character, and expanding the central region in a second direction and a third direction according to a first expansion step size; stopping further expansion when the edge of the central region expanded along the second direction meets a preset edge search termination condition, and determining the current edge position of the expanded central region in the second direction as the coordinate position of the text character in the second direction; and further expanding the central region in the third direction... If the edge of the central region satisfies the preset edge search termination condition, the expansion stops, and the edge position of the currently expanded central region in the third direction is determined as the coordinate position of the text character in the third direction; the coordinate position is used as the layout position parameter of the text character; the receptive field corresponding to each character is intermediate data in the neural network execution process; the preset edge search termination condition includes at least one of the following: Condition 1: The pixels of the currently expanded central region disappear at the edge of the corresponding expansion direction; Condition 2: The receptive field of the slice to which the currently expanded central region belongs in the corresponding expansion direction is identified as covering a character that is a transition character or a segmentation character; Condition 3: The receptive field of the slice to which the currently expanded central region belongs in the corresponding expansion direction is identified as covering another character that is different from the text character; The text character is set in the editable text format according to the layout position parameters of the text character.

2. The method according to claim 1, characterized in that, The step of identifying the image to be identified based on the set receptive field and determining the character information in the image to be identified includes: Determine the character coverage area of ​​the receptive field in the image to be recognized; If it is determined that the receptive field covers a complete character in the image to be recognized, the text character of the complete character covered by the receptive field is recognized, and the layout position parameters of the text character are also recognized. If it is determined that the receptive field does not cover the complete character in the image to be recognized, the receptive field is moved along the first direction by a first moving step.

3. The method according to claim 2, characterized in that, After determining that the receptive field covers a complete character in the image to be recognized, the method further includes: Determine whether the complete character covered by the receptive field contains unrecognized characters; If it is determined that the complete character covered by the receptive field contains unrecognized characters, the text characters of the unrecognized characters in the complete character covered by the receptive field are identified, as well as the layout position parameters of the text characters; If it is determined that the complete character covered by the receptive field does not contain unrecognized characters, the receptive field is moved along the first direction by a first movement step.

4. The method according to claim 2, characterized in that, The step of identifying text characters of complete characters covered by the receptive field in the image to be identified, when it is determined that the receptive field covers the complete character, includes: Determine the pixel region in the image to be recognized that is covered by the complete character of the receptive field; Based on the pixels of the pixel region corresponding to the complete character, the corresponding text character is generated.

5. The method according to claim 1, characterized in that, The process of identifying the image to be identified and determining the character information in the image to be identified includes: Extract character features from the image to be identified and convert the character features into a character feature sequence; Identify and decode the character feature sequence to obtain the text character corresponding to the character feature sequence; Calculate the coordinate position of the text character and use the coordinate position as the layout position parameter of the text character.

6. The method according to any one of claims 1-5, characterized in that, The image to be identified includes: one character, multiple characters, or a whole text; wherein, when the image to be identified includes a whole text, the whole text includes multiple adjacent characters.

7. The method according to claim 6, characterized in that, The process of identifying the image to be identified and determining the character information in the image to be identified includes: The method involves identifying the image to be identified and determining the character information of the entire text segment in the image; or, identifying the image to be identified and determining the character information of a preset number of characters in the image; or, identifying the image to be identified and determining the character information of characters of a preset language type in the image.

8. A character recognition device, characterized in that, include: The acquisition module is used to acquire the image to be recognized; The recognition module is used to identify the image to be recognized using a neural network based on a set receptive field, determine character information in the image to be recognized, and obtain text characters and their layout position parameters based on the character information. Obtaining the layout position parameters of the text characters includes: during text character recognition, determining that the text character corresponds to one or more slices in the receptive field, defining the slice corresponding to the text character in the receptive field as the central region of the text character, and expanding the central region in a second direction and a third direction according to a first expansion step size; stopping further expansion when the edge of the central region expanded along the second direction meets a preset edge search termination condition, and determining the edge position of the currently expanded central region in the second direction as the coordinate position of the text character in the second direction; and expanding along the third direction... If the edge of the central region of the expansion meets the preset edge search termination condition, the expansion stops, and the edge position of the currently expanded central region in the third direction is determined as the coordinate position of the text character in the third direction; the coordinate position is used as the layout position parameter of the text character; the receptive field corresponding to each character is the intermediate data of the neural network execution process; the preset edge search termination condition includes at least one of the following: Condition 1: The pixels of the currently expanded central region disappear at the edge of the corresponding expansion direction; Condition 2: The receptive field of the slice to which the currently expanded central region belongs in the corresponding expansion direction is identified as covering a character that is a transition character or a segmentation character; Condition 3: The receptive field of the slice to which the currently expanded central region belongs in the corresponding expansion direction is identified as covering another character that is different from the text character; The conversion module is used to set the text characters in an editable text format according to the layout position parameters of the text characters.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program for character recognition, and the processor is configured to read and run the computer program for character recognition to perform the character recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the character recognition method according to any one of claims 1 to 7 when it runs.

Citation Information

Patent Citations

  • Digital image processing system applied to bill image character recognition and method

    CN104112128A

  • Verification code identification method and apparatus, computer device and computer storage medium

    CN107688809A

  • Image processing system and an image processing method

    US20190303702A1