A text correction method and system based on bounding boxes

By detecting text information in the image and forming a bounding box, detecting and rotating the bounding box to correct the text direction, the problem that the prior art cannot handle different rotation angles and inconsistent rotation angles of multiple segments of text is solved, and effective correction of various directions and angle texts in the image is achieved.

CN113903038BActive Publication Date: 2025-06-10GUANGDONG KAMFU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111191591.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2025-06-10
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

The prior art cannot effectively correct the situation where the different rotation angles of text information in images and the rotation angles of multiple segments of text are inconsistent.

Method used

The text correction method based on the bounding box is adopted, by receiving images, detecting text information, forming a rectangular bounding box, detecting the angle between the normal vector of the bounding box and the vertical direction, and rotating the bounding box according to the included angle, identifying and correcting the direction of the text information.

Benefits of technology

This method can be applied to correct text information in various directions and angles in the image, facilitate text recognition, and solves the limitation that traditional methods can only deal with the consistent rotation angle of all text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903038B_ABST
    Figure CN113903038B_ABST
Patent Text Reader

Abstract

The present invention discloses a text correction method and system based on bounding boxes, which includes receiving an image, detecting text information in the image by using the text detection algorithm DB, and forming rectangular bounding boxes around each text information; taking the normal vector along the long side of the bounding box and detecting the angle between the normal vector and the vertical direction; comparing the angle with a preset angle value, and if the angle is greater than the preset angle value, rotating the bounding box to within the preset angle value according to the angle; identifying the direction of the text information in the bounding box, and if the direction of the text information is upside down, rotating the upside-down bounding box by 180 degrees, and extracting the text information containing the keyword information in the bounding box according to the pre-entered keyword information. It is used to correct text information in various directions and at various angles in the image information, and solves the problem that the traditional text correction method can only correct the situation where all text rotation angles are the same.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a text correction method and system based on a bounding box. Background Art

[0002] When acquiring images, the pictures obtained by taking photos usually have the situation of non - upright surface images, which affects the accuracy of image recognition. Therefore, it is necessary to correct the pictures taken.

[0003] The existing correction methods only consider the case where all text has the same rotation angle and can only handle rotation angles of 90 degrees, 180 degrees, and 270 degrees. When encountering other rotation angles or the situation where the rotation angles of multiple text materials are inconsistent, the existing methods cannot correct the text simultaneously. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention provides a text correction method and system based on a bounding box.

[0005] In order to achieve the above - mentioned purpose, the technical solutions adopted by the present invention are as follows:

[0006] According to the first - aspect embodiment of the present invention, a text correction method based on a bounding box is provided, including: S10. Receive an image, use the text detection algorithm DB to detect text information in the image, and form a rectangular bounding box around each piece of text information; S20. Take the normal vector along the long side of the bounding box and detect the angle between the normal vector and the vertical direction; S30. Compare the angle with a preset angle value. If the angle is greater than the preset angle value, rotate the bounding box to within the preset angle value according to the angle; S40. Identify the direction of the text information in the bounding box. If the direction of the text information is upside - down, rotate the upside - down bounding box by 180 degrees, and extract the text information containing the keyword information in the bounding box according to the pre - entered keyword information.

[0007] According to some embodiments of the present invention, before step S10, there is also step S1, and step S1 includes inputting or deleting keyword information and key layout.

[0008] According to some embodiments of the present invention, step S10 further includes, if some keyword information is lost, searching for the un - located keyword information through regular expressions and the located keyword information.

[0009] According to some embodiments of the present invention, in step S30, if the angle between the normal vector and the vertical direction is less than or equal to the preset angle value, directly enter step S40, and the preset angle value is set to 5 degrees.

[0010] According to some embodiments of the present invention, step S30 further includes taking the center point of the bounding box as a fixed point, rotating the bounding box, and re-entering step S20 to detect the angle between the normal vector and the vertical direction.

[0011] According to some embodiments of the present invention, after step S30, there is further step S31, and step S31 includes detecting whether there is an overlap between two adjacent bounding boxes. If there is no overlap between the bounding boxes, directly enter step S40. If there is an overlap between the two bounding boxes, cut the second bounding box according to the edge of the first bounding box to obtain two bounding boxes after the first cut, and cut the first bounding box according to the edge of the second bounding box to obtain two bounding boxes after the second cut, and select the two bounding boxes with the highest confidence.

[0012] According to the embodiments of the second aspect of the present invention, there is provided a text correction system based on a bounding box, including: a positioning module, configured to receive an image, detect text information in the image by using the text detection algorithm DB, and form a rectangular bounding box around each piece of text information; a detection module, configured to take a normal vector along the long side of the bounding box and detect the angle between the normal vector and the vertical direction; a correction module, configured to compare the angle with a preset angle value. If the angle is greater than the preset angle value, rotate the bounding box to within the preset angle value according to the angle; an extraction module, configured to identify the direction of the text information in the bounding box. If the direction of the text information is upside down, rotate the bounding box with the upside-down direction by 180 degrees, and extract the text information containing the keyword information in the bounding box according to the pre-entered keyword information.

[0013] According to some embodiments of the present invention, there is further an input module, and the input module includes inputting or deleting keyword information and key layout.

[0014] According to some embodiments of the present invention, the positioning module further includes, if some keyword information is lost, searching for the un-located keyword information through regular expressions and the located keyword information.

[0015] According to some embodiments of the present invention, there is further a cutting module, and the cutting module includes detecting whether there is an overlap between two adjacent bounding boxes. If there is an overlap between the two bounding boxes, cut the second bounding box according to the edge of the first bounding box to obtain two bounding boxes after the first cut, and cut the first bounding box according to the edge of the second bounding box to obtain two bounding boxes after the second cut, and select the two bounding boxes with the highest confidence.

[0016] A method and system for text correction based on a bounding box according to an embodiment of the present invention have at least the following beneficial effects: By forming a bounding box around each text information and rotating and correcting each bounding box with incorrect text orientation, it can be applied to correct text information in various directions and at various angles in the image information, facilitating text recognition, and solving the problem that the traditional text correction method can only correct the case where all text rotation angles are the same.

[0017] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0019] Figure 1 is a flowchart of the present invention;

[0020] Figure 2 is a schematic diagram of an image that needs text correction of the present invention;

[0021] Figure 3 is a schematic diagram after forming the bounding box and the normal vector of the present invention;

[0022] Figure 4 is a schematic diagram of bounding box correction of the present invention;

[0023] Figure 5 is a schematic diagram of bounding box overlapping and cutting of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further explains the specific implementation methods of the present invention in conjunction with the drawings.

[0025] The technical solutions in the embodiments of the present invention will be described completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] The present invention provides a method for text correction based on a bounding box, as Figure 1The following is a flowchart of the present invention. According to some embodiments, in step S10, an image is received, and text information is detected in the image using the text detection algorithm DB, and a rectangular bounding box is formed around each piece of text information; in step S20, a normal vector is taken along the long side of the bounding box, and the angle between the normal vector and the vertical direction is detected; in step S30, the angle is compared with a preset angle value. If the angle is greater than the preset angle value, the bounding box is rotated to within the preset angle value according to the angle; in step S40, the direction of the text information within the bounding box is recognized. If the direction of the text information is upside down, the upside-down bounding box is rotated 180 degrees, and the text information containing the keyword information within the bounding box is extracted according to the pre-entered keyword information.

[0027] Based on the above embodiments, in some embodiments, as Figures 2 - 4 shown, the image information contains multiple pieces of text information, and the directions of each piece of text information are all different. Before correction, first, the position of the text information is found in the image using the text detection algorithm DB, and a bounding box is formed around the text information. The bounding box is a rectangular bounding box, and a normal vector is taken on one side of the long side of the bounding box, and it is ensured that the direction of the normal vector is perpendicular to the long side of the bounding box. The direction of the bounding box is associated with the text direction. It should be noted that when the direction of the normal vector of the bounding box is vertically upward, the text information within the bounding box is in the horizontal direction. Therefore, this solution needs to detect the angle between the direction of the normal vector of the bounding box and the vertical direction. When the angle is greater than the preset angle value set in advance, the bounding box is rotated according to the angle. As Figure 2 shown, there are bounding boxes in multiple directions, and each bounding box rotates with the angle between its own normal vector direction and the vertical direction as the rotation angle. After rotation, the bounding box may be upside down. When the bounding box is upside down, the angle between the normal vector and the vertical direction is 0. After rotation, the text directions within all bounding boxes are recognized. If the direction of the text information is recognized as upside down, the corresponding bounding box is rotated 180 degrees for correction, and after correction, the text information is extracted. Traditional text correction and image correction methods can only rotate the entire image or all the text in the image by the same angle for correction. However, when applying the traditional correction method to an image such as Figure 2 shown, not all text can be corrected. Therefore, the present invention aims to solve the deficiencies of the prior art and provides a method capable of performing rotation correction of different angles and different directions on multiple texts.

[0028] According to some embodiments, before step S10, there is also step S1, and step S1 includes entering or deleting keyword information and key layout.

[0029] Based on the above embodiments, when there is unimportant text information in the image, such as manually writing unimportant text information on a copy, after taking an image of the copy with a document camera, the keyword information can be more accurately located according to the keyword information and the key layout, and the key information can be found based on the keyword information, where the keyword information is "name", "gender", "birth", "year", "month", "day", etc., and the key information is the specific person names and place names, etc., before and after the keyword information.

[0030] According to some embodiments, step S10 further includes, if some keyword information is lost, using regular expressions and the located keyword information to find the unlocated keyword information.

[0031] Based on the above embodiments, when holding an ID card information and taking a photo of the ID card with a document camera, some keyword information may be covered by the hand. Use regular expressions and the located keyword information to find the position of the unlocated keyword information and supplement the keyword information.

[0032] According to some embodiments, in step S30, if the included angle between the normal vector and the vertical direction is less than or equal to a preset angle value, directly enter step S40, and the preset angle value is set to 5 degrees.

[0033] Based on the above embodiments, the preset angle value is set to 5 degrees. When the clockwise or counterclockwise offset angle of the image is less than or equal to 5 degrees, it does not affect the recognition of the text. Therefore, when the clockwise or counterclockwise offset of the image is less than or equal to 5 degrees, image rotation is not required.

[0034] According to some embodiments, step S30 further includes, taking the center point of the bounding box as a fixed point, rotating the bounding box, and re-entering step S20 to detect the included angle between the normal vector and the vertical direction.

[0035] Based on the above embodiments, the bounding box is a rectangular bounding box, and the diagonal focus of the bounding box is its center point. Rotate the bounding box with the center point of the bounding box as a fixed point. After the rotation is completed, re-enter step S20 to detect whether the included angle between the normal vector of the bounding box and the vertical direction is less than or equal to the preset angle value. If the included angle is less than or equal to the preset angle value, enter the next step; otherwise, repeat steps S20 → step S30 → step S20 until the included angle is less than or equal to the preset angle value.

[0036] According to some embodiments, step S30 also includes step S31, and step S31 includes detecting whether there is overlap between two adjacent bounding boxes. If there is no overlap between the bounding boxes, directly entering step S40. If there is overlap between the two bounding boxes, the second bounding box is clipped according to the edge of the first bounding box to obtain two bounding boxes clipped for the first time, and the first bounding box is clipped according to the edge of the second bounding box to obtain two bounding boxes clipped for the second time, and the two bounding boxes with the highest confidence are selected.

[0037] Based on the above embodiments, Figure 4 , Figure 5 As shown in , when there is overlap between two bounding boxes, keep the first bounding box to clip the second bounding box and keep the second bounding box to clip the first bounding box, and keep the two bounding boxes with the highest confidence. Figure 5 As shown in the figure, the text in the four bounding boxes after clipping is not lost, the confidence is high, and the text in the bounding box can be recognized. When the text in the bounding box is missing, move the two bounding boxes to the ends and clip again, or take the case where the bounding box is not missing, such as Figure 5 The bounding box 1 and the bounding box 4 shown are complete no matter how they are cropped, and their confidence is the highest. The bounding box 1 and the bounding box 4 with the highest confidence are selected for text information extraction and recognition.

[0038] On the basis of the above embodiments, this embodiment provides a text correction system based on a bounding box. According to some embodiments, a positioning module is used to receive an image, detect text information in the image using a text detection algorithm DB, and form a rectangular bounding box around each of the text information; a detection module is used to take a normal vector from the long side of the bounding box and detect the angle between the normal vector and the vertical direction; a correction module is used to compare the angle with a preset angle value, and if the angle is greater than the preset angle value, rotate the bounding box to within the preset angle value according to the angle; an extraction module is used to identify the direction of the text information in the bounding box, and if the direction of the text information is inverted, rotate the inverted bounding box 180 degrees, and extract the text information containing the keyword information in the bounding box according to the pre-entered keyword information.

[0039] Based on the above embodiments, in some embodiments, such as Figures 2 - 4As shown, the image information contains multiple text information, and the directions of each text information are different. Before correction, first use the text detection algorithm DB to find the position of the text information in the image, and form a bounding box around the text information. The bounding box is a rectangular bounding box, and a normal vector is taken on one side of the long side of the bounding box, and it is ensured that the direction of the normal vector is perpendicular to the long side of the bounding box. The direction of the bounding box is associated with the text direction. It should be noted that when the direction of the normal vector of the bounding box is vertically upward, the text information in the bounding box is in the horizontal direction. Therefore, this solution needs to detect the angle between the direction of the normal vector of the bounding box and the vertical direction. When the angle is greater than a preset angle value, the bounding box is rotated according to the angle. As Figure 2 shown, there are bounding boxes in multiple directions, and each bounding box is rotated with the angle between its own normal vector direction and the vertical direction as the rotation angle. After rotation, the bounding box may be upside down. When the bounding box is upside down, the angle between the normal vector and the vertical direction is 0. After rotation, the text directions in all bounding boxes are recognized. If the direction of the text information is recognized as upside down, the corresponding bounding box is rotated 180 degrees for correction. After correction, the text information is extracted. Traditional text correction and image correction methods can only rotate the entire image or all the text in the image by the same angle for correction, and using the traditional correction method for an image such as Figure 2 shown, not all text can be corrected. Therefore, the present invention aims to solve the defects of the prior art and provide a method capable of performing rotation correction on multiple texts at different angles and in different directions.

[0040] According to some embodiments, it further includes an input module, and the input module includes inputting or deleting keyword information and key layout.

[0041] Based on the above embodiments, when there is unimportant text information in the image, such as manually writing unimportant text information on a copy and then taking a picture of the copy image with a document camera, according to the keyword information and key layout, the keyword information can be more accurately located, and the key information can be found according to the keyword information. The keyword information includes "name", "gender", "birth", "year", "month", "day", etc., and the key information is the specific person names and place names, etc. before and after the keyword information.

[0042] According to some embodiments, the positioning module further includes, if some keyword information is lost, using regular expressions and the located keyword information to find the un-located keyword information.

[0043] Based on the above embodiments, when holding an identity card information and taking a photo of the ID card with a document camera, some keyword information may be covered by the hand. Using regular expressions and the located keyword information to find the position of the un-located keyword information and supplement the keyword information.

[0044] According to some embodiments, it further includes a clipping module. The clipping module includes detecting whether there is an overlap between two adjacent bounding boxes. If there is an overlap between the two bounding boxes, the second bounding box is clipped according to the edge of the first bounding box to obtain two bounding boxes after the first clipping, and the first bounding box is clipped according to the edge of the second bounding box to obtain two bounding boxes after the second clipping, and two bounding boxes with the highest confidence are selected.

[0045] Based on the above embodiments, as Figure 4 , Figure 5 shown, when there is an overlap between two bounding boxes, the first bounding box is maintained to clip the second bounding box and the second bounding box is maintained to clip the first bounding box, and two bounding boxes with the highest confidence are retained. As Figure 5 shown, the text in the four bounding boxes after clipping is not lost, the confidence is high, and the text in the bounding box can be recognized. When there is missing text in the bounding box, the two bounding boxes are moved to both ends and clipped again, or the case where there is no missing text in the bounding box is taken, such as Figure 5 the bounding box 1 and the bounding box 4 shown. No matter how they are clipped, the bounding box 1 and the bounding box 4 in both clipping methods are complete, and their confidence is the highest. The bounding box 1 and the bounding box 4 with the highest confidence are taken for text information extraction and recognition.

[0046] For those skilled in the art, it is obvious that the present invention is not limited to the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the basic features of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting.

Claims

1. A text correction system applying a text correction method based on bounding boxes, characterized in that, it includes: A positioning module, which is used to receive an image, detect text information in the image by using the text detection algorithm DB, and form a rectangular bounding box around each piece of text information; the positioning module also includes, if some keyword information is lost, searching for the unpositioned keyword information through regular expressions and the positioned keyword information; A detection module, which is used to take the normal vector along the long side of the bounding box and detect the angle between the normal vector and the vertical direction; A correction module, which is used to compare the angle with a preset angle value. If the angle is greater than the preset angle value, rotate the bounding box to within the preset angle value according to the angle; An extraction module, which is used to identify the direction of the text information in the bounding box. If the direction of the text information is upside down, rotate the upside-down bounding box by 180 degrees, and extract the text information containing keyword information in the bounding box according to the pre-entered keyword information; It also includes an input module, and the input module includes inputting or deleting keyword information and key layout; it also includes a clipping module, and the clipping module includes detecting whether there is an overlap between two adjacent bounding boxes. If there is an overlap between the two bounding boxes, clip the second bounding box according to the edge of the first bounding box to obtain two bounding boxes after the first clipping, and clip the first bounding box according to the edge of the second bounding box to obtain two bounding boxes after the second clipping, and select the two bounding boxes with the highest confidence; The text correction method includes: S10. Receive an image, detect text information in the image by using the text detection algorithm DB, and form a rectangular bounding box around each piece of text information; S20. Take the normal vector along the long side of the bounding box and detect the angle between the normal vector and the vertical direction; S30. Compare the angle with a preset angle value. If the angle is greater than the preset angle value, rotate the bounding box to within the preset angle value according to the angle; S40. Identify the direction of the text information in the bounding box. If the direction of the text information is upside down, rotate the upside-down bounding box by 180 degrees, and extract the text information containing keyword information in the bounding box according to the pre-entered keyword information; In step S30, if the angle between the normal vector and the vertical direction is less than or equal to the preset angle value, directly enter step S40, and the preset angle value is set to 5 degrees; Step S30 also includes taking the center point of the bounding box as a fixed point, rotating the bounding box, and re-entering step S20 to detect the angle between the normal vector and the vertical direction; After step S30, there is also step S31, and step S31 includes detecting whether there is an overlap between two adjacent bounding boxes. If there is no overlap between the bounding boxes, directly enter step S40. If there is an overlap between the two bounding boxes, clip the second bounding box according to the edge of the first bounding box to obtain two bounding boxes after the first clipping, and clip the first bounding box according to the edge of the second bounding box to obtain two bounding boxes after the second clipping, and select the two bounding boxes with the highest confidence; Before the step S10, there is also a step S1, and the step S1 includes inputting or deleting keyword information and key formatting; the step S10 also includes that if some keyword information is lost, the un-located keyword information is found through regular expressions and the located keyword information.

Citation Information

Patent Citations

  • Image processing method and system, electronic equipment and storage medium

    CN113420762A

  • Character recognition method, character recognition device and storage medium

    WO2021146937A1