Text content extraction and recognition method based on vertical merging of line text boxes
By pointing fingers to the topic and combining image preprocessing and detection algorithms, vertical text boxes merge is realized, solving the problems of complex operation, time-consuming and low accuracy in the prior art, and improving the recognition efficiency and accuracy.
Patent Information
- Application Number
- CN202210499796.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The prior art is complex, time-consuming and accurate when taking pictures and entering questions into software to obtain questions, and the existing vertical merging methods fail to effectively select merge lines based on text location and semantic information.
Pointing to the topic by fingers under the device camera, using image preprocessing, text detection, hand area and key point detection, combined with text box information and finger node position, the text box is merged vertically and the content of the topic is output.
It improves the accuracy and speed of text recognition, simplifies the operation process, reduces server resource requirements, and realizes efficient vertical text box merging.
Smart Images

Figure CN114821620B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and text recognition, and particularly relates to a text content extraction and recognition method based on vertical merging of line text boxes. Background Art
[0002] Currently, one related product after another born based on AI technology is entering ordinary people's homes. They are practical, easy to use, and can meet the personalized needs of users, empowering various learning scenarios in all aspects. For example, by taking a photo of an entire piece of paper including the questions, and then cropping the position where the questions are located to obtain the semantic information of the questions. The process of taking a photo to search for questions is to first use OCR (Optical Character Recognition) to process and recognize the questions in the picture into text, and then compare the user's question text with the question bank in the platform database to find the top 5 most similar ones. During the process of processing and recognizing the questions in the picture into text, the operation is relatively complex and time-consuming; or in some related educational software, it is necessary to input the content of the questions into the software to obtain the relevant questions. The above methods obtain relevant questions by inputting, and the accuracy is not high, and it is easy to find incorrect questions; if we want to improve the above problems such as operation complexity, time consumption, and accuracy, it is necessary to involve relevant technologies such as OCR technology, text content extraction and merging in computer vision. OCR refers to the process in which an electronic device (such as a scanner or digital camera) checks the characters on paper, determines their shapes by detecting dark and bright patterns, and then uses character recognition methods to translate the shapes into computer text descriptions. Text box merging is to vertically merge several line text boxes into the same text box.
[0003] Chinese Patent No. CN113963342A provides a line merging method based on text box position and character information. This method combines the characteristics of Chinese characters and the characteristics of the detection algorithm, uses the size information and position information of the text box to merge lines, restores the text line information, has a fast processing speed, and the merged line text conforms to the text information of the original picture, improving the overall recognition accuracy; however, this patented technology only horizontally selects the text belonging to one line with a single box, achieving a horizontal merge, and does not achieve a vertical merge.
[0004] The Chinese patent with the publication number CN113850208A provides a method, device, equipment and medium for structuring picture information. It uses a text recognition network and a text detection network to detect and recognize text in a picture, so as to obtain the first minimum bounding rectangle of each detected text box and the corresponding text information. The text boxes obtained by the text detection network are sorted in a preset order, and the text information of all text boxes is merged according to the sorting result to obtain the text content in the picture, and the information of the target label is extracted from the text content by using regular rules. Although this patented technology performs vertical merging, it merges all lines without selecting certain lines for merging according to text position information and semantic information. Summary of the Invention
[0005] In view of the above, the present invention provides a method for extracting and recognizing text content based on vertical merging of line text boxes. By pointing at the question to be viewed with a finger under the device camera and uploading it to the server in the form of a picture, all the text related to this question can be framed with a rectangular box through a simple algorithm, realizing vertical merging of one line, and presenting the text in the rectangular box.
[0006] A method for extracting and recognizing text content based on vertical merging of line text boxes includes the following steps:
[0007] (1) For the text image of a test paper or exercise book, first preprocess the image, and then use the existing text detection algorithm to extract and recognize all text boxes and their information in the image;
[0008] (2) Use the existing object detection model to frame the hand area pointing to the question in the form of a rectangular box, and then detect the position information of each key node of the finger within the hand area;
[0009] (3) Utilize the text box information and the position information of the finger key nodes to vertically merge all text boxes pointed to by the finger belonging to the same question into a rectangular box and extract it;
[0010] (4) Use the existing text recognition method to recognize and obtain the content in the merged rectangular box, and this content is the text description of the question pointed to by the finger.
[0011] Further, the preprocessing of the image in step (1) includes image perspective transformation and mean filter denoising processing, where the transformation matrix used for image perspective transformation is automatically adjusted according to the height and angle of the captured picture; using this method can improve the accuracy of subsequent text detection and text recognition.
[0012] Further, in the step (1), a text detection algorithm based on PaddleOCR is used to identify and extract the text boxes in the image. On the basis of using the Paddle pre-trained model, training is carried out using the text image dataset of test papers and exercise books, which can effectively extract the text boxes including text, punctuation marks, and underlines, effectively making up for the situations of incorrect detection and recognition failure of Paddle, and improving the accuracy of model detection and recognition.
[0013] Further, the text box information extracted and recognized in the step (1) includes the positions of the four vertices of the text box, the text content within the text box, and the confidence level.
[0014] Further, in the step (2), the YOLOv5 model is used to frame the hand area pointing to the question in the form of a rectangular box, and at the same time, the position information of each key node of the finger is detected by using skeleton detection.
[0015] Further, the specific implementation process of the step (3) is as follows:
[0016] 3.1 Initially screen the text boxes that meet the conditions according to the position information of the finger tip nodes;
[0017] 3.2 Find all the text boxes of the pointed question from the set of text boxes saved by the initial screening according to the position information and content information of the text boxes;
[0018] 3.3 Merge all the text boxes of the pointed question into a rectangular box.
[0019] Further, the specific implementation method of the step 3.1 is as follows: First, select the text boxes that meet the condition of x_left ≤ x1 ≤ x_right from all the text boxes in the image; then find the text boxes that meet the condition of y_left < y1 among the selected text boxes as the initial screening result; where x_left is the x-axis coordinate value of the upper left vertex of the text box, x_right is the x-axis coordinate value of the upper right vertex of the text box, y_left is the y-axis coordinate value of the upper left vertex of the text box, and x1 and y1 are the x-axis coordinate value and y-axis coordinate value of the finger tip node respectively.
[0020] Further, the specific implementation of step 3.2 is as follows: First, find the text box closest to the fingertip node from the preliminarily screened and saved text box set and denote it as T1, and T1 is the text box of the referred question. Determine whether there is a question number at the beginning of T1. If there is, stop the search; if not, find the text box closest to T1 above it from the text box set and denote it as T2, and determine whether T1 and T2 satisfy the following relationship. If not, discard T2 and stop the search; if satisfied, determine T2 as the text box of the referred question, and then determine whether there is a question number at the beginning of T2. If there is, stop the search; if not, continue to search upward according to the above method until all text boxes in the text box set are judged.
[0021] x3 ≤ x2 ≤ x4 and frame_height ≥ frame_distance
[0022] Where: x2 is the x-axis coordinate value of the center point of T1, x3 is the x-axis coordinate value of the upper left vertex of T2, x4 is the x-axis coordinate value of the upper right vertex of T2, frame_height is the frame height of T2, and frame_distance is the frame distance between T2 and T1.
[0023] Further, the specific implementation of step 3.3 is as follows: For all text boxes of the referred question, find the maximum x-axis coordinate value max_x, the maximum y-axis coordinate value max_y, the minimum x-axis coordinate value min_x, and the minimum y-axis coordinate value min_y among the four vertices of these text boxes. Then establish a rectangular box with the lower left vertex coordinate as (min_x, min_y), the lower right vertex coordinate as (max_x, min_y), the upper right vertex coordinate as (max_x, max_y), and the upper left vertex coordinate as (min_x, max_y). The text enclosed by this rectangular box is all the content of the question pointed by the finger.
[0024] Based on the position information, semantic information, and finger coordinate information of the text box, the present invention uses a simple and efficient algorithm to implement the merging of multiple related text boxes and outputs the content in the merged text box. This algorithm is simple and efficient, solving the problem of insufficient server resources. At the same time, the present invention uses object detection and hand detection recognition to find the coordinates of the finger key points. After using the existing model and subsequent training, the accuracy is improved and the speed becomes faster, which can better cooperate with the text box merging algorithm. Brief Description of the Drawings
[0025] Figure 1 It is a flow chart of the text content extraction and recognition method of the present invention.
[0026] Figure 2 It is the effect diagram after image preprocessing.
[0027] Figure 3The image result obtained after text detection.
[0028] Figure 4 The image result obtained after hand region detection.
[0029] Figure 5 The image result obtained after hand key point detection.
[0030] Figure 6 The example diagram of the text box after preliminary screening.
[0031] Figure 7 The example diagram of the text box after further screening.
[0032] Figure 8 The example diagram of the visualization of the title box.
[0033] Figure 9 The example diagram of the final result obtained by text recognition. Specific implementation manner
[0034] In order to describe the present invention more specifically, the technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0035] As Figure 1 shown, the text content extraction and recognition method of the present invention based on the vertical merging of line text boxes, by merging relevant line text boxes into a rectangular box and outputting the text in the merged rectangular box (for example, using a rectangular box to select the lines belonging to the same question in the student exercise book and discarding the text boxes that do not belong to this question), the specific implementation includes the following steps:
[0036] (1) Image preprocessing.
[0037] The present invention uses image perspective transformation, brightness enhancement, and mean filter denoising algorithms to process the picture. Among them, the perspective transformation matrix M is automatically adjusted according to the height and angle of the captured picture. Using this method, the accuracy of subsequent text detection and text recognition can be improved.
[0038] The effect of the preprocessed image is as Figure 2 shown.
[0039] (2) Text detection.
[0040] Perform text detection on the preprocessed image. In this embodiment, a text detection algorithm based on paddleocr is used to select the text in the image. Based on the paddle pre-trained model, a dataset suitable for this scenario is used for training, and the parameters of the text detection model are fine-tuned, which can effectively extract text boxes including text, punctuation marks, and underlines, effectively making up for the situations of paddle detection recognition errors and recognition failures, and improving the accuracy of model detection and recognition.
[0041] Through the trained model, all the text box texts in the picture can be recognized, and the positions, contents, and confidences of the four vertices of the text box are returned. The image result obtained after text detection is as Figure 3 shown.
[0042] (3) Hand region detection.
[0043] Use an object detection algorithm to detect the position of the hand in the original image. In this embodiment, the yolov5 model is used to select the hand region with a rectangular box. Based on the object detection pre-trained model, pictures with hand regions are used for training, and the parameters of the object detection model are fine-tuned. This approach can effectively extract the hand region, facilitating the next step of finger key point detection.
[0044] The image result obtained after hand region detection is as Figure 4 shown.
[0045] (4) Finger key point detection.
[0046] In the hand region obtained in step (3) of the present invention, finger bone key point detection can obtain the position information of each key point of the hand, including the coordinate information (x, y) of the finger tip point. Then, according to the perspective transformation matrix M, the (x, y) coordinates are transformed into (x1, y1). This operation can accurately find the finger position information for positioning the text box position below. The image result obtained after hand key point detection is as Figure 5 shown.
[0047] The formula for transforming the (x, y) coordinates into (x1, y1) is as follows:
[0048] x1 = (M[0][0] * p[0] + M[0][1] * p[1] + M[0][2]) / ((M[2][0] * p[0] + M[2][1] * p[1] + M[2][2])) y1 = (M[1][0] * p[0] + M[1][1] * p[1] + M[1][2]) / ((M[2][0] * p[0] + M[2][1] * p[1] + M[2][2]))
[0049] (5) Text box screening module.
[0050] The present invention screens text boxes using the position information, content information of text boxes and finger coordinates. The algorithm is simple and can effectively screen out the questions that are closest to the finger and belong to the same question according to the position information and semantic information of the text boxes, so as to vertically merge the text boxes belonging to the same question below. The specific implementation process is as follows
[0051] Step 1: According to finger positioning, initially screen the line text boxes that meet the conditions.
[0052] According to the finger coordinates x1 obtained by the four finger key point detections, find the text boxes with the left upper corner x value less than x1 and the right upper corner x value greater than x1; according to the finger coordinates y1, among all the text boxes that meet the x range, find the text boxes with the left upper corner y value less than y1 of the text boxes, and initially screen the text boxes.
[0053] The judgment criterion is the line text boxes where x_left <= x1 <= x_right and y_left < y1, where x_left is the left upper corner x value of the text box, x_right is the right upper corner x value of the text box, and y_left is the left upper corner y value of the text box.
[0054] An example of the text boxes after the initial screening is as Figure 6 shown.
[0055] Step 2: Further screen the line text boxes according to the position information and semantic information of the text boxes.
[0056] Among the text boxes that meet the above initial screening process, since the position information of the text boxes is stored in the list in the order from top to bottom and from left to right, the last text box in the list is the text box closest to the finger, and this text box is one line of the question pointed by the finger. Calculate the center point in the x direction of this text box (denoted as x2), and check whether the beginning of this line of text is a question number. If not, find the text box closest to this line of text box upward, calculate the frame height frame_height of this text box, and calculate the frame distance frame_distance between the text box framed by this line of text and the text box framed by the line closest to the finger. If it satisfies that the frame height of this line of text box is greater than the frame distance between the two text boxes, that is, frame_height > frame_distance, and x2 is between the leftmost x value and the rightmost x value of this line of text box, then save this line, and this line is one line of the question to be found; if the condition is not met, discard it and stop traversing the remaining text boxes; if it is judged that the beginning of the text box is a question number, stop looking upward, then regard the text boxes framed by the lines that have been found as the whole question, and save the information of the text boxes framed by the lines that have been found, including the information of the four vertices, text information and confidence.
[0057] The judgment criterion is to start searching upward from the text box closest to the finger coordinate point. Find the text box that meets the conditions of x3 <= x1 <= x4 and frame_height >= frame_distance, retain it and continue the upward repetition operation. Otherwise, discard it and stop traversing the remaining text boxes. Here, x2 is the x value of the center point of the previous text box, x3 and x4 correspond to the x value of the upper left corner and the x value of the upper right corner of the next text box respectively, frame_height is the height of the next text box, and frame_distance is the distance between the previous text box and the next text box.
[0058] According to the above, judge each text box upward in turn until all text boxes are judged or the traversal stops. Save the position information of the text boxes that meet the conditions. An example of the further filtered text boxes is as Figure 7 shown; The algorithm of this operation process is simple, does not require too many GPU server resources, has a fast response speed, and is more in line with the scenario used in the present invention.
[0059] (6) Text box merging module.
[0060] In the text boxes that meet the conditions in step (5) of the present invention, find the maximum x value, the maximum y value, the minimum x value, and the minimum y value among the four vertices of these text boxes, and set them as max_x, max_y, min_x, and min_y in turn. Then use these four values to form four coordinate points, which are (min_x, min_y), (max_x, min_y), (max_x, max_y), and (min_x, max_y) in turn. These four points are used as the lower left corner, lower right corner, upper right corner, and upper left corner of the rectangle frame. The text enclosed by this rectangle frame is all the content of the question to be found; By doing so, the algorithm is simple, which can effectively shorten the operation time on the server, has a fast response speed, and high accuracy.
[0061] An example of the visualization of the question box is as Figure 8 shown.
[0062] (7) Text recognition.
[0063] The present invention uses the text recognition algorithm with fine-tuned parameters for the final screening result of the text boxes obtained in step (6) to obtain the content in the text boxes. This content is all the text descriptions of the question pointed by the finger. Using the text recognition algorithm with fine-tuned parameters can effectively improve the accuracy of text recognition in this scenario, including the recognition of underlines and punctuation marks.
[0064] An example of the final result of text recognition is as Figure 9 shown.
[0065] The above description of the embodiments is provided to enable those of ordinary skill in the art to understand and apply the present invention. It is obvious that those who are familiar with the technology in this field can easily make various modifications to the above embodiments, and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the protection scope of the present invention.
Claims
1. A text content extraction and recognition method based on vertical merging of line text boxes, comprising the following steps: (1) For text images of test papers and exercise books, first preprocess the images, and then use existing text detection algorithms to extract and recognize all text boxes and their information in the images; (2) Use an existing object detection model to frame the hand area pointing to the question in the form of a rectangular box, and then detect the position information of each key node of the finger within the hand area; (3) Utilize the text box information and the position information of the finger key nodes to vertically merge all text boxes pointed to by the finger belonging to the same question into a rectangular box for extraction. The specific implementation process is as follows: 3.1 According to the position information of the finger tip node, initially screen the text boxes that meet the conditions. Specifically: First, select the text boxes that satisfy x_left ≤ x1 ≤ x_right from all text boxes in the image; then find the text boxes that satisfy y_left < y1 among the selected text boxes as the initial screening result; where x_left is the x-axis coordinate value of the upper left vertex of the text box, x_right is the x-axis coordinate value of the upper right vertex of the text box, y_left is the x-axis coordinate value of the upper left vertex of the text box, and x1 and y1 are the x-axis coordinate value and y-axis coordinate value of the finger tip node respectively; 3.2 According to the position information and content information of the text boxes, find all text boxes of the question pointed to from the set of text boxes saved by the initial screening. Specifically: First, find the text box closest to the finger tip node from the set of text boxes saved by the initial screening and denote it as T1, and T1 is the text box of the question pointed to. Judge whether there is a question number at the beginning of T1. If there is, stop the search; if not, find the text box closest to T1 above it from the set of text boxes and denote it as T2. Judge whether T1 and T2 satisfy the following relationship. If not, discard T2 and stop the search; if satisfied, determine that T2 is the text box of the question pointed to, and then judge whether there is a question number at the beginning of T2. If there is, stop the search. If not, continue to search upward according to the above method until all text boxes in the set of text boxes are judged; x3 ≤ x2 ≤ x4 and frame_height ≥ frame_distance Wherein: x2 is the x-axis coordinate value of the center point of T1, x3 is the x-axis coordinate value of the upper left vertex of T2, x4 is the x-axis coordinate value of the upper right vertex of T2, frame_height is the frame height of T2, and frame_distance is the frame distance between T2 and T1; 3.3 Merge all the text boxes of the indicated question into a rectangular box. Specifically: for all the text boxes of the indicated question, find the maximum x-axis coordinate value max_x, the maximum y-axis coordinate value max_y, the minimum x-axis coordinate value min_x, and the minimum y-axis coordinate value min_y among the four vertices of these text boxes. Then establish a rectangular box with the lower left vertex coordinate as (min_x, min_y), the lower right vertex coordinate as (max_x, min_y), the upper right vertex coordinate as (max_x, max_y), and the upper left vertex coordinate as (min_x, max_y). The text enclosed by this rectangular box is all the content of the question pointed by the finger. (4) Use the existing text recognition method to recognize and obtain the content in the merged rectangular box, which is the text description of the question pointed by the finger.
2. The text content extraction and recognition method according to claim 1, characterized in that: The preprocessing of the image in step (1) includes image perspective transformation and mean filter denoising processing. Among them, the transformation matrix used for image perspective transformation is automatically adjusted according to the height and angle of the captured picture.
3. The text content extraction and recognition method according to claim 1, characterized in that: In step (1), the text detection algorithm based on PaddleOCR is used to recognize and extract the text boxes in the image. On the basis of using the Paddle pre-trained model, it is trained with the text image dataset of test papers and exercise books, and the text boxes including text, punctuation marks, and underlines can be effectively extracted.
4. The text content extraction and recognition method according to claim 1, characterized in that: The information of the text boxes extracted and recognized in step (1) includes the positions of the four vertices of the text box, the text content inside the text box, and the confidence level.
5. The text content extraction and recognition method according to claim 1, characterized in that: In step (2), the YOLOv5 model is used to frame the hand area pointing to the question in the form of a rectangular box, and at the same time, the position information of each key node of the finger is detected by using skeleton detection.
Citation Information
Patent Citations
Picture information structuring method and device, equipment and medium
CN113850208A
A textbox position and character information-based row merging method
CN113963342A
Subtitle extraction method and device based on artificial intelligence, equipment and storage medium
CN114359942A