Character recognition method and device, program product, electronic equipment and storage medium
By using the character box detection model and the character trace detection model in the automated correction system, combined with the interchange and comparison screening, the problem that the existing system cannot accurately identify the manual correction trace is solved, and the accurate recognition of the target characters on written materials is achieved.
Patent Information
- Application Number
- CN202510239347.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The existing automated correction system cannot accurately identify traces of manual correction on written materials.
By inputting the image to be identified into the character box detection model and the character trace detection model, the positions of the character box and the target characters are determined respectively, and the effective target characters are filtered and recognized according to the interleaving ratio.
Accurate identification of target characters on written materials, especially manual correction traces, is achieved, and the accuracy of the automated correction system is improved.
Smart Images

Figure CN120182980A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of character recognition, and particularly to a character recognition method, device, program product, electronic device, and storage medium. Background Art
[0002] With the development of information technology and artificial intelligence, automated marking systems have gradually become a research hotspot. These systems use technologies such as optical character recognition (OCR), natural language processing (NLP), and machine learning to automatically analyze and mark the written materials submitted by students.
[0003] However, there is still a large amount of manual marking in the fields of education and training. For example, teachers and assessors usually need to mark students' homework, exam papers, and other forms of written materials, that is, add marking traces indicating right or wrong, scores, suggestions, etc. The automated marking systems in the related technologies are still unable to accurately recognize the manual marking traces on the written materials. Summary of the Invention
[0004] To overcome the problems in the related technologies, embodiments of the present disclosure provide a character recognition method, device, program product, electronic device, and storage medium to solve the defects in the related technologies.
[0005] According to a first aspect of the embodiments of the present disclosure, a character recognition method is provided. The method includes:
[0006] Inputting an image to be recognized into a character box detection model to obtain the positions of at least one character box output by the character box detection model, where the character box is used to fill in target characters;
[0007] Inputting the image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model;
[0008] Determining valid target characters among the at least one target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character.
[0009] In a possible embodiment of the present disclosure, the determining valid target characters among the at least one target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character includes:
[0010] Determining the intersection over union (IoU) of each character box and each target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character;
[0011] Traverse each target character in sequence, and when traversing the target character: if the intersection over union (IoU) between the character box with the largest IoU among the unlocked character boxes and the target character is greater than the first threshold, determine that the target character is a valid target character, and determine that the character box with the largest IoU with the target character is locked.
[0012] In a possible embodiment of the present disclosure, determining the IoU between each character box in the at least one character box and each target character in the at least one target character according to the positions of each character box and each target character includes:
[0013] Determine the overlapping area between the character box and the target character according to the positions of the character box and the target character;
[0014] If the area of the character box is greater than the area of the target character, determine the ratio between the overlapping area and the area of the character box as the IoU between the character box and the target character;
[0015] If the area of the character box is not greater than the area of the target character, determine the ratio between the overlapping area and the area of the target character as the IoU between the character box and the target character.
[0016] In a possible embodiment of the present disclosure, after traversing all the target characters, the method further includes:
[0017] For each unlocked character box, crop the image to be recognized according to the position of the character box to obtain a local image of the character box, and input the local image of the character box into the character trace cutout recognition model to obtain a first recognition result output by the character trace cutout recognition model, where the first recognition result includes whether there is a target character or not in the character box.
[0018] In a possible embodiment of the present disclosure, after traversing all the target characters, the method further includes:
[0019] For each invalid target character, crop the image to be recognized according to the position of the target character to obtain a local image of the target character, and input the local image of the target character into the character trace cutout recognition model to obtain a second recognition result output by the character trace cutout recognition model, where the second recognition result includes whether the target character is valid or not.
[0020] In a possible embodiment of the present disclosure, cropping the image to be recognized according to the position of the target character to obtain a local image of the target character includes:
[0021] If the difference between the position of the target character and the average position of the character frames is less than a second threshold, crop the image to be recognized according to the position of the target character to obtain a local image of the target character, where the average position of the character frames is the average of the positions of all character frames.
[0022] In a possible embodiment of the present disclosure, after traversing all target characters, the method further includes:
[0023] Crop the image to be recognized according to the average position of the character frames to obtain a character frame area image, where the average position of the character frames is the average of the positions of all character frames, and the character frame area image is an area image that covers the average position of the character frames and conforms to the distribution mode of the character frames;
[0024] Perform binarization processing on the character frame area image to determine at least one connected domain within the character area image;
[0025] For each connected domain that has no intersection with the character frame and the target character, input the connected domain into a character trace cut image recognition model to obtain a third recognition result output by the character trace cut image recognition model, where the third recognition result includes that there is a target character within the connected domain or there is no target character within the connected domain.
[0026] In a possible embodiment of the present disclosure, the step of inputting the connected domain into a character trace cut image recognition model to obtain a third recognition result output by the character trace cut image recognition model includes:
[0027] If the area of the connected domain is greater than a third threshold and less than a fourth threshold, input the connected domain into a character trace cut image recognition model to obtain a third recognition result output by the character trace cut image recognition model.
[0028] In a possible embodiment of the present disclosure, the step of inputting the image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model includes:
[0029] Input the image to be recognized into a character trace detection model to obtain the positions and types of at least one target character output by the character trace detection model.
[0030] In a possible embodiment of the present disclosure, the target character is a corrected character.
[0031] According to a second aspect of the embodiments of the present disclosure, there is provided a character recognition device, the device includes:
[0032] A character box recognition module, configured to input an image to be recognized into a character box detection model to obtain the positions of at least one character box output by the character box detection model, where the character box is used to fill in target characters;
[0033] A character recognition module, configured to input an image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model;
[0034] A determination module, configured to determine valid target characters from the at least one target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character.
[0035] In a possible embodiment of the present disclosure, the determination module is configured to:
[0036] Determine the intersection over union (IoU) of each character box and each target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character;
[0037] Traverse each target character in sequence, and when traversing the target character: if the IoU of the character box with the largest IoU among the unlocked character boxes and the target character is greater than a first threshold, determine that the target character is a valid target character, and determine that the character box with the largest IoU with the target character is locked.
[0038] In a possible embodiment of the present disclosure, when the determination module is configured to determine the IoU of each character box and each target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character, it is configured to:
[0039] Determine the overlapping area of the character box and the target character according to the position of the character box and the position of the target character;
[0040] If the area of the character box is greater than the area of the target character, determine the ratio between the overlapping area and the area of the character box as the IoU of the character box and the target character;
[0041] If the area of the character box is not greater than the area of the target character, determine the ratio between the overlapping area and the area of the target character as the IoU of the character box and the target character.
[0042] In a possible embodiment of the present disclosure, the apparatus further includes a first supplement module, configured to:
[0043] After traversing all target characters, for each unlocked character box, crop the image to be recognized according to the position of the character box to obtain a local image of the character box, and input the local image of the character box into the character trace cropping recognition model to obtain a first recognition result output by the character trace cropping recognition model, where the first recognition result includes that there is a target character or no target character in the character box.
[0044] In a possible embodiment of the present disclosure, the device further includes a first supplement module for:
[0045] After traversing all target characters, crop the image to be recognized according to the position of the target character to obtain a local image of the target character, and input the local image of the target character into the character trace cropping recognition model to obtain a second recognition result output by the character trace cropping recognition model, where the second recognition result includes that the target character is valid or invalid.
[0046] In a possible embodiment of the present disclosure, the cropping the image to be recognized according to the position of the target character to obtain a local image of the target character includes:
[0047] If the difference between the position of the target character and the average position of the character boxes is less than a second threshold, crop the image to be recognized according to the position of the target character to obtain a local image of the target character, where the average position of the character boxes is the average of the positions of all character boxes.
[0048] In a possible embodiment of the present disclosure, the device further includes a first supplement module for:
[0049] After traversing all target characters, crop the image to be recognized according to the average position of the character boxes to obtain a character box area image, where the average position of the character boxes is the average of the positions of all character boxes, and the character box area image is an area image covering the average position of the character boxes and conforming to the distribution mode of the character boxes;
[0050] Perform binarization processing on the character box area image to determine at least one connected domain in the character area image;
[0051] For each connected domain that has no intersection with both the character box and the target character, input the connected domain into the character trace cropping recognition model to obtain a third recognition result output by the character trace cropping recognition model, where the third recognition result includes that there is a target character or no target character in the connected domain.
[0052] In a possible embodiment of the present disclosure, when the third supplementary module is used to input the connected component into the character trace cutting image recognition model to obtain a third recognition result output by the character trace cutting image recognition model, it is used for:
[0053] If the area of the connected component is greater than a third threshold and less than a fourth threshold, the connected component is input into the character trace cutting image recognition model to obtain a third recognition result output by the character trace cutting image recognition model.
[0054] In a possible embodiment of the present disclosure, the character recognition module is used for:
[0055] Input the image to be recognized into the character trace detection model to obtain the positions and types of at least one target character output by the character trace detection model.
[0056] In a possible embodiment of the present disclosure, the target character is a corrected character.
[0057] According to a third aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer programs / instructions, which when executed by a processor, implement the steps of the method described in the first aspect.
[0058] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, where the electronic device includes a memory and a processor, the memory is used to store computer instructions that can run on the processor, and the processor is used to implement the method described in the first aspect when executing the computer instructions.
[0059] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.
[0060] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0061] The character recognition method provided by the embodiments of the present disclosure first inputs the image to be recognized into a character box detection model to obtain the positions of at least one character box output by the character box detection model, then inputs the image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model, and finally determines valid target characters from the at least one target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character. Since the character box for filling in the target character and the target character are recognized separately, and a secondary judgment is made on the recognized target character based on the recognition results of both, valid target characters can be accurately recognized. If the target character is a manual correction trace on a written material with a preset character box, this method can accurately recognize the correction trace on the written material. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.
[0063] Figure 1 is a schematic diagram of a terminal device shown in an exemplary embodiment of the present disclosure;
[0064] Figure 2 is a flowchart of a character recognition method shown in an exemplary embodiment of the present disclosure;
[0065] Figure 3 is a schematic diagram of an image to be processed shown in an exemplary embodiment of the present disclosure;
[0066] Figure 4 is a flowchart of a method for screening valid target characters shown in an exemplary embodiment of the present disclosure;
[0067] Figure 5 is a schematic structural diagram of a character recognition device shown in an exemplary embodiment of the present disclosure;
[0068] Figure 6 is a block diagram of the structure of an electronic device shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0070] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0071] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0072] In the related art, an automated marking system can detect and classify different types of marks by training a convolutional neural network (CNN); or use a recurrent neural network (RNN) to understand the semantic information in handwritten annotations. However, these methods often result in low recognition accuracy due to factors such as insufficient datasets, the diversity of handwritten scripts, and complex background interference.
[0073] Based on this, in a first aspect, at least one embodiment of this disclosure provides a character recognition method that can accurately recognize target characters on a written material with a preset character box, such as a manual marking trace on a written material with a preset character box.
[0074] For example, this method can be applied to Figure 1 the cloud server connected to terminal devices 1, 2, and 3 in the appendix, or any one of terminal devices 1, 2, 3, 4, 5, and 6; if applied to the cloud server, the cloud server can receive character data sent by terminal devices 1, 2, or 3 connected thereto and recognize it using this method, and if applied to a terminal device, the terminal device can recognize local collected character data (such as a character image) using this method.
[0075] Please refer to the appendix Figure 2 , which exemplarily shows the flow of the character recognition method, including steps S201 to S203.
[0076] In step S201, the image to be recognized is input into a character box detection model to obtain the positions of at least one character box output by the character box detection model, where the character box is used to fill in the target character.
[0077] Please refer to the appendix Figure 3, the image to be recognized is an image of a written material, on which character frames are pre-labeled (e.g., printed). The character frames are used to fill in target characters, such as for filling in handwritten target characters. For example, the written material is a student's test paper, and a character frame for marking can be added to the right side of each question in the test paper. This character frame is used for the teacher to fill in marking characters, such as characters indicating right or wrong (e.g., a tick for right, a cross and a circle for wrong), characters indicating scores, characters indicating suggestions, etc. Preferably, the character frames in the student's test paper can be divided into multiple areas by horizontal or vertical dividing lines to add marking characters in different marking links. Preferably, the character frames on the written material can have position consistency in one or more directions. For example, the character frames in the student's test paper can be close to the right edge of the test paper in the horizontal direction of the test paper.
[0078] By pre-labeling character frames on the written material, it is possible to assist in the recognition of target characters and improve the convenience and accuracy of character recognition. Taking the teacher's manual marking as an example, usually when a teacher marks a student's homework or test paper, the marking is done at the position where the student answers. Such a marking method will cause the marking traces to appear at any position in the image. The sizes of the marking symbols marked by different teachers are also different, and the marking backgrounds are also various, which will cause great difficulties in the recognition of marking traces; while the present disclosure makes the positions of the teacher's marking traces neat and uniform, and the sizes are not very different, and the backgrounds are the same, so that the recognition of marking traces can be easier and more accurate.
[0079] The present disclosure places no restrictions on the specific type of the character frame detection model. For example, the character frame detection model can be formed by training with the yolo5 model.
[0080] Exemplarily, the image to be recognized can be pre-processed first, such as scaled to a specific size, image data standardized, etc., and then the image to be recognized is input into the character frame detection model.
[0081] Among them, the character frame can be rectangular, and the position of the character frame can be the position coordinates of two diagonal vertices of the character frame, such as the position coordinates of the upper left corner and the lower right corner.
[0082] In step S202, the image to be recognized is input into the character trace detection model to obtain the positions of at least one target character output by the character trace detection model. For example, the target character is a marking character.
[0083] The present disclosure places no restrictions on the specific type of the character trace detection model. For example, the character trace detection model can be formed by training with the yolo5 model.
[0084] Exemplarily, the image to be recognized can be preprocessed first, such as scaled to a specific size, image data standardized, etc., and then the image to be recognized is input into the character trace detection model.
[0085] In addition, the character trace detection model can also output the type of the target character. For example, if the target character is a correction trace, the type of the target character can be the type of the correction trace, such as right, wrong, etc.
[0086] Among them, the position of the target character can be the position coordinates of the two diagonal vertices of the rectangular area frame where the target character is located, such as the position coordinates of the upper left corner and the lower right corner.
[0087] In step S203, according to the position of each character box in the at least one character box and the position of each target character in the at least one target character, valid target characters are determined in the at least one target character.
[0088] Exemplarily, it can be determined whether each target character is a valid target character in the manner shown in the appendix Figure 4 including steps S401 to S402.
[0089] In step S401, according to the position of each character box in the at least one character box and the position of each target character in the at least one target character, the intersection over union (IoU) of each character box and each target character is determined.
[0090] For example:
[0091] First, the overlapping area between the character box and the target character is determined according to the position of the character box and the position of the target character; for example, the length and width of the character box are determined based on the position coordinates of the upper left corner and the lower right corner of the character box, and then the area of the character box is determined; for example, the length and width of the rectangular area frame are determined based on the position coordinates of the upper left corner and the lower right corner of the rectangular area frame where the target character is located, and then the area of the rectangular area frame is determined as the area of the target character.
[0092] Next, if the area of the character box is greater than the area of the target character, the ratio between the overlapping area and the area of the target character box is determined as the intersection over union of the character box and the target character; if the area of the character box is not greater than the area of the target character, the ratio between the overlapping area and the area of the character box is determined as the intersection over union of the character box and the target character. That is, the ratio between the overlapping area and the area of the smaller one of the character box and the target area is determined as the intersection over union of the character box and the target character.
[0093] In step S402, each target character is traversed in sequence, and when traversing the target character: if the intersection-over-union ratio of the character box with the largest intersection-over-union ratio with the target character in the unlocked character boxes is greater than the first threshold, it is determined that the target character is a valid target character, and the character box with the largest intersection-over-union ratio with the target character is determined to be locked.
[0094] Among them, a locked character box refers to a character box in which a target character has been determined to exist, that is, when determining that a certain target character is a valid target character, the character box where the target character is located is synchronously determined.
[0095] For example, a marking report can be further generated according to the type of valid target characters. For example, the marking report includes marking items such as the question ID, the correctness of the question, and the score.
[0096] The character recognition method provided by the embodiments of the present disclosure first inputs the image to be recognized into a character box detection model to obtain the positions of at least one character box output by the character box detection model, then inputs the image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model, and finally determines valid target characters among the at least one target characters according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character. Since the character box for filling in the target character and the target character are respectively recognized, and a secondary judgment is made on the recognized target characters based on the recognition results of the two to determine valid target characters, the target characters in the image to be recognized can be accurately recognized. If the target character is a manual marking trace on a written material with a preset character box, this method can accurately recognize the marking trace on the written material, accurately recognize the marking trace of the teacher on the paper material, which can not only help the automated marking system better understand the evaluation intention, but also provide more detailed and personalized feedback for students.
[0097] In some embodiments of the present disclosure, after traversing all target characters, some target characters are determined to be valid target characters, and the remaining target characters can be called invalid target characters, and some character boxes are determined to be locked, and the remaining character boxes are unlocked.
[0098] Based on the above situation, exemplarily, the method further includes: for each invalid target character, cropping the image to be recognized according to the position of the target character to obtain a local image of the target character, and inputting the local image of the target character into a character trace cutout recognition model to obtain a second recognition result output by the character trace cutout recognition model, where the second recognition result includes whether the target character is valid or invalid.
[0099] For example, if the difference between the position of the target character and the average position of the character frames is less than a second threshold, the image to be recognized is cropped according to the position of the target character to obtain a local image of the target character, where the average position of the character frames is the average of the positions of all character frames. For example, the average position of the character frames refers to the average of the central coordinates of all character frames. If the difference between the central coordinate of the target character and the average position of the character frames is less than the second threshold, the image to be recognized is cropped according to the position of the target character to obtain a local image of the target character. Preferably, if the character frames on the written material have position consistency in the horizontal direction, the difference between the horizontal coordinate value of the central coordinate of the target character and the horizontal coordinate value of the average position of the character frames is used to represent the coordinate difference between the two; if the character frames on the written material have position consistency in the vertical direction, the difference between the vertical coordinate value of the central coordinate of the target character and the vertical coordinate value of the average position of the character frames is used to represent the coordinate difference between the two. For example, the character frames in a student's test paper can be close to the right edge of the test paper in the horizontal direction of the test paper, and the difference between the horizontal coordinate value of the central coordinate of the target character and the horizontal coordinate value of the average position of the character frames can be used to represent the coordinate difference between the two.
[0100] The present disclosure does not limit the specific type of the character trace cropping and recognition model. For example, the character trace cropping and recognition model can be formed by training using the mobilenetV3 small model. The training data of the character trace cropping and recognition model includes samples in which there are target characters in the image but the character trace detection model cannot detect them, and samples in which there are character frames in the image but the character frame detection model cannot detect them.
[0101] For example, the local image of the target character can be preprocessed first, such as scaling to a specific size, performing image data normalization, etc., and then the local image of the target character is input into the character trace cropping and recognition model.
[0102] In addition, the character trace cropping and recognition model can also output the type of the target character. For example, if the target character is a correction trace, the type of the target character can be the type of the correction trace, such as right, wrong, etc. It can be understood that the character trace cropping and recognition model can output the type of the target character or the type of the non-target character. If the type of the target character such as right, wrong, etc. is output, it is determined that the target character in the local image is a valid target character. If the type of the non-target character is output, it is determined that the target character in the local image is an invalid target character.
[0103] This example performs secondary recognition on the target character through the character trace cut-out recognition model to accurately confirm its validity. Thus, in the case where the target character is recognized successfully but the character box where it is located is not recognized successfully, the judgment result of the validity of the target character can be corrected to avoid missing the recognition of valid target characters.
[0104] Based on the above situation, exemplarily, the method further includes: for each unlocked character box, cropping the image to be recognized according to the position of the character box to obtain a partial image of the character box, and inputting the partial image of the character box into the character trace cut-out recognition model to obtain a first recognition result output by the character trace cut-out recognition model, where the first recognition result includes whether there is a target character in the character box or not.
[0105] For example, the partial image can be cropped according to the boundary of the character box, or the partial image can be cropped after extending a certain length in each direction along the boundary of the character box.
[0106] The present disclosure places no restrictions on the specific type of the character trace cut-out recognition model. For example, the character trace cut-out recognition model can be formed by training with the mobilenetV3 small model.
[0107] For example, the partial image of the character box can be preprocessed first, such as scaling to a specific size, standardizing the image data, etc., and then the partial image of the character box is input into the character trace cut-out recognition model.
[0108] In addition, the character trace cut-out recognition model can also output the type of the target character. For example, if the target character is a correction trace, the type of the target character can be the type of the correction trace, such as right, wrong, etc. It can be understood that the character trace cut-out recognition model can output the type of the target character or the type of the non-target character. If the type of the target character such as right, wrong, etc. is output, it is determined that there is a target character in the partial image of the character box. If the type of the non-target character is output, it is determined that there is no target character in the partial image of the character box.
[0109] This example performs secondary recognition on the character box through the character trace cut-out recognition model to accurately confirm whether there is a target character therein. Thus, in the case where the character box is recognized successfully but the target character therein is not recognized successfully, it can accurately recognize whether there is a target character therein, avoiding missing the recognition of valid target characters.
[0110] Based on the above situation, exemplarily, the method further includes the following steps:
[0111] First, crop the image to be recognized according to the average position of the character frames to obtain a character frame area image, where the average position of the character frames is the average of the positions of all character frames, and the character frame area image is an area image that covers the average position of the character frames and conforms to the distribution pattern of the character frames.
[0112] Among them, the average position of the character frames may include the average of the central coordinates of all character frames, the average of the widths (horizontal lengths) of all character frames, and the average of the heights (vertical lengths) of all character frames.
[0113] Among them, the distribution pattern of the character frames refers to the positional consistency of the character frames in one or more directions.
[0114] For example, the character frames have positional consistency in the horizontal direction. Let avgX be the average of the horizontal axis coordinate values in the central coordinates of all character frames, avgW be the average of the widths (horizontal lengths) of all character frames, and H be the height of the image to be processed. The character frame area image is an area image with avgX - avgW as the left boundary, avgX + avgW as the right boundary, 0 as the upper boundary, and H - 1 as the lower boundary.
[0115] Next, perform binarization processing on the character frame area image to determine at least one connected component within the character area image.
[0116] For example, use the OTSU method to perform binarization processing on the character frame area image. Among them, a connected component refers to an area composed of multiple adjacent pixels with a pixel value of 1, or an area composed of multiple adjacent pixels with a pixel value of 0.
[0117] Finally, for each connected component that has no intersection with the character frames and the target characters, input the connected component into the character trace cutting and recognition model to obtain a third recognition result output by the character trace cutting and recognition model, where the third recognition result includes that there is a target character within the connected component or there is no target character within the connected component.
[0118] Among them, a connected component that has no intersection with the character frames and the target characters refers to a connected component that has no intersection with each character frame and has no intersection with each target character.
[0119] For example, if the area of the connected component is greater than a third threshold and less than a fourth threshold, the connected component is input into the character trace cutout recognition model to obtain a third recognition result output by the character trace cutout recognition model. For example, the third threshold is 1 / 2 of the average area of all character frames, and the fourth threshold is 2 times the average area of all character frames. In this way, connected components with a small area can be filtered out as artifacts, and at the same time, connected components with an overly large area can be excluded as invalid connected components, thereby reducing the number of connected components input into the character trace cutout recognition model and improving the efficiency of the secondary recognition of the target character in this example.
[0120] The present disclosure does not limit the specific type of the character trace cutout recognition model. For example, the character trace cutout recognition model can be formed by training a mobilenetV3 small model.
[0121] For example, the connected component can be preprocessed first, such as scaled to a specific size, image data normalized, etc., and then the connected component is input into the character trace cutout recognition model.
[0122] In addition, the character trace cutout recognition model can also output the type of the target character. For example, if the target character is a correction trace, the type of the target character can be the type of the correction trace, such as right, wrong, etc. It can be understood that the character trace cutout recognition model can output the type of the target character or the type of the non-target character. If the type of the target character such as right, wrong, etc. is output, it is determined that there is a valid target character in the connected component. If the type of the non-target character is output, it is determined that there is no valid target character in the connected component.
[0123] In this example, the character trace cutout recognition model performs secondary recognition on the image region of the distributed character frames to accurately confirm whether there is a missed target character therein. Thus, the target character can be accurately recognized additionally when both the character frame and the target character therein are not recognized successfully, avoiding missing valid target characters.
[0124] According to a second aspect of the embodiments of the present disclosure, a character recognition device is provided. Please refer to the attached Figure 5 , the device includes:
[0125] A character frame recognition module 501, configured to input an image to be recognized into a character frame detection model to obtain the positions of at least one character frame output by the character frame detection model, where the character frame is used to fill in the target character;
[0126] A character recognition module 502, configured to input the image to be recognized into a character trace detection model to obtain the positions of at least one target character output by the character trace detection model;
[0127] A determination module 503, configured to determine valid target characters among the at least one target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character.
[0128] In a possible embodiment of the present disclosure, the determination module is configured to:
[0129] Determine the intersection over union (IoU) of each character box and each target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character;
[0130] Traverse each target character in sequence, and when traversing the target character: if the IoU of the character box with the largest IoU among the unlocked character boxes and the target character is greater than a first threshold, determine that the target character is a valid target character, and determine that the character box with the largest IoU with the target character is locked.
[0131] In a possible embodiment of the present disclosure, when the determination module is configured to determine the IoU of each character box and each target character according to the positions of each character box in the at least one character box and the positions of each target character in the at least one target character, it is configured to:
[0132] Determine the overlapping area between the character box and the target character according to the position of the character box and the position of the target character;
[0133] If the area of the character box is greater than the area of the target character, determine the ratio between the overlapping area and the area of the character box as the IoU of the character box and the target character;
[0134] If the area of the character box is not greater than the area of the target character, determine the ratio between the overlapping area and the area of the target character as the IoU of the character box and the target character.
[0135] In a possible embodiment of the present disclosure, the apparatus further includes a first supplementary module, configured to:
[0136] After traversing all target characters, for each unlocked character box, crop the image to be recognized according to the position of the character box to obtain a local image of the character box, and input the local image of the character box into a character trace cutout recognition model to obtain a first recognition result output by the character trace cutout recognition model, where the first recognition result includes whether there is a target character or no target character in the character box.
[0137] In a possible embodiment of the present disclosure, the apparatus further includes a first supplementary module, configured to:
[0138] After traversing all target characters, crop the image to be recognized according to the positions of the target characters to obtain local images of the target characters, and input the local images of the target characters into a character trace cropping recognition model to obtain a second recognition result output by the character trace cropping recognition model, where the second recognition result includes whether the target characters are valid or invalid.
[0139] In a possible embodiment of the present disclosure, the cropping the image to be recognized according to the positions of the target characters to obtain local images of the target characters includes:
[0140] If the difference between the position of the target character and the average position of the character frames is less than a second threshold, crop the image to be recognized according to the position of the target character to obtain a local image of the target character, where the average position of the character frames is the average of the positions of all character frames.
[0141] In a possible embodiment of the present disclosure, the device further includes a first supplementary module for:
[0142] After traversing all target characters, crop the image to be recognized according to the average position of the character frames to obtain a character frame area image, where the average position of the character frames is the average of the positions of all character frames, and the character frame area image is an area image covering the average position of the character frames and conforming to the character frame distribution pattern;
[0143] Perform binarization processing on the character frame area image to determine at least one connected component within the character area image;
[0144] For each connected component that has no intersection with both the character frame and the target character, input the connected component into a character trace cropping recognition model to obtain a third recognition result output by the character trace cropping recognition model, where the third recognition result includes whether there is a target character within the connected component or there is no target character within the connected component.
[0145] In a possible embodiment of the present disclosure, when the third supplementary module is used to input the connected component into a character trace cropping recognition model to obtain a third recognition result output by the character trace cropping recognition model, it is used for:
[0146] If the area of the connected component is greater than a third threshold and less than a fourth threshold, input the connected component into a character trace cropping recognition model to obtain a third recognition result output by the character trace cropping recognition model.
[0147] In a possible embodiment of the present disclosure, the character recognition module is used for:
[0148] Input the image to be recognized into the character trace detection model to obtain the positions and types of at least one target character output by the character trace detection model.
[0149] In a possible embodiment of the present disclosure, the target character is a correction character.
[0150] Regarding the device in the above embodiment, the specific manners in which each model performs operations have been described in detail in the embodiments of the method in the first aspect, and will not be elaborated herein.
[0151] According to a third aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program / instructions, which when executed by a processor implement the steps of the method described in the first aspect.
[0152] According to a fourth aspect of the embodiments of the present disclosure, please refer to the appendix Figure 6 , which exemplarily shows a block diagram of an electronic device. For example, the device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0153] Refer to Figure 6 , the device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0154] The processing component 602 generally controls the overall operation of the device 600, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 602 may include one or more models to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia model to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0155] The memory 604 is configured to store various types of data to support the operation of the device 600. Examples of such data include instructions for any application or method operating on the device 600, contact data, phone book data, messages, pictures, videos, and the like. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0156] The power component 606 provides power to the various components of the device 600. The power component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 600.
[0157] The multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0158] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.
[0159] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface model, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0160] The sensor assembly 614 includes one or more sensors for providing a status assessment of various aspects of the device 600. For example, the sensor assembly 614 can detect the on / off state of the device 600, the relative positioning of components, such as the display and keypad of the device 600. The sensor assembly 614 can also detect a change in the position of the device 600 or a component of the device 600, the presence or absence of user contact with the device 600, the orientation or acceleration / deceleration of the device 600, and the temperature change of the device 600. The sensor assembly 614 can also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0161] The communication component 616 is configured to facilitate communication between the device 600 and other devices in a wired or wireless manner. The device 600 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0162] In an exemplary embodiment, the device 600 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the character recognition method of the above-mentioned electronic device.
[0163] In a fifth aspect, in an exemplary embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium including instructions, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the device 600 to complete the character recognition method of the above-mentioned electronic device. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0164] Other embodiments of the present disclosure will be readily apparent to those skilled in the art in view of the specification and practice of the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0165] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A character recognition method, characterized in that: The method comprises: Inputting the image to be recognized into a character box detection model to obtain the position of at least one character box output by the character box detection model, wherein the character box is used to fill in the target character; Inputting the image to be recognized into a character trace detection model to obtain the position of at least one target character output by the character trace detection model; According to the position of each character box in the at least one character box and the position of each target character in the at least one target character, a valid target character is determined in the at least one target character.
2. The character recognition method according to claim 1, characterized in that: The determining a valid target character in the at least one target character according to the position of each character box in the at least one character box and the position of each target character in the at least one target character comprises: Determine an intersection-over-unit ratio between each character box and each target character according to the position of each character box in the at least one character box and the position of each target character in the at least one target character; Each target character is traversed in turn, and when traversing the target character: if the character box with the largest intersection-and-union ratio with the target character among the unlocked character boxes has an intersection-and-union ratio with the target character greater than a first threshold, then the target character is determined to be a valid target character, and it is determined that the character box with the largest intersection-and-union ratio with the target character is locked.
3. The character recognition method according to claim 2, characterized in that: The determining, according to the position of each character box in the at least one character box and the position of each target character in the at least one target character, an intersection-over-union ratio of each character box and each target character, comprises: Determining an overlapping area between the character frame and the target character according to the position of the character frame and the position of the target character; If the area of the character frame is larger than the area of the target character, the ratio between the overlapping area and the area of the character frame is determined as the intersection-over-union ratio of the character frame and the target character; If the area of the character frame is not larger than the area of the target character, the ratio between the overlapping area and the area of the target character is determined as the intersection-over-union ratio of the character frame and the target character.
4. The character recognition method according to claim 2, characterized in that: After traversing all target characters, the method further includes: For each unlocked character box, the image to be recognized is cropped according to the position of the character box to obtain a partial image of the character box, and the partial image of the character box is input into the character trace cutting image recognition model to obtain a first recognition result output by the character trace cutting image recognition model, wherein the first recognition result includes the presence or absence of the target character in the character box.
5. The character recognition method according to claim 2, characterized in that: After traversing all target characters, the method further includes: For each invalid target character, the image to be recognized is cropped according to the position of the target character to obtain a partial image of the target character, and the partial image of the target character is input into a character trace cut image recognition model to obtain a second recognition result output by the character trace cut image recognition model, wherein the second recognition result includes whether the target character is valid or invalid.
6. The character recognition method according to claim 5, characterized in that: The step of cropping the image to be recognized according to the position of the target character to obtain a partial image of the target character includes: If the difference between the position of the target character and the average position of the character frame is less than a second threshold, the image to be recognized is cropped according to the position of the target character to obtain a partial image of the target character, wherein the average position of the character frame is the average value of the positions of all character frames.
7. The character recognition method according to claim 2, characterized in that: After traversing all target characters, the method further includes: Cropping the image to be recognized according to the average position of the character frames to obtain a character frame area image, wherein the average position of the character frames is the average value of the positions of all the character frames, and the character frame area image is an area image covering the average position of the character frames and conforming to the character frame distribution method; Binarizing the character frame region image to determine at least one connected domain in the character region image; For each connected domain that has no intersection with the character box and the target character, the connected domain is input into the character trace graph cutting recognition model to obtain a third recognition result output by the character trace graph cutting recognition model, wherein the third recognition result includes the existence of the target character in the connected domain or the absence of the target character in the connected domain.
8. The character recognition method according to claim 7, characterized in that: The step of inputting the connected domain into a character trace graph cut recognition model to obtain a third recognition result output by the character trace graph cut recognition model includes: If the area of the connected domain is greater than the third threshold and less than the fourth threshold, the connected domain is input into the character trace graph cut recognition model to obtain a third recognition result output by the character trace graph cut recognition model.
9. The character recognition method according to claim 1, characterized in that: The step of inputting the image to be recognized into the character trace detection model to obtain the position of at least one target character output by the character trace detection model comprises: The image to be recognized is input into the character trace detection model to obtain the position and type of at least one target character output by the character trace detection model.
10. The character recognition method according to claim 1, characterized in that: The target character is a correction character.
11. A character recognition device, characterized in that: The device comprises: A character box recognition module, used for inputting the image to be recognized into a character box detection model to obtain the position of at least one character box output by the character box detection model, wherein the character box is used to fill in the target character; A character recognition module, used for inputting the image to be recognized into a character trace detection model to obtain the position of at least one target character output by the character trace detection model; The determination module is used to determine a valid target character in the at least one target character according to the position of each character box in the at least one character box and the position of each target character in the at least one target character.
12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store computer instructions executable on the processor, and the processor is used to implement the method according to any one of claims 1 to 10 when executing the computer instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.