Character recognition method and device
Patent Information
- Application Number
- CN202380010979.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the accuracy of fingertip character positioning and recognition is low, especially when text is bent and large text is incomplete.
By obtaining screenshots of the area where the user's fingertips are located, text detection is performed to obtain the area text box, and combining single-word detection and character positioning information, the target character pointed to by the user's fingertips are determined. The method includes steps such as screenshot acquisition, text detection, matching, single-word detection and character recognition, and uses a deep learning network for image processing and recognition.
It significantly improves the accuracy of fingertip character positioning and recognition, can effectively handle curved text and large text recognition, and improves the efficiency and accuracy of reading assistive technology.
Smart Images

Figure CN120129935A_ABST
Abstract
Description
Character recognition method and device Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a character recognition method and device. Background Art
[0002] Reading assistive technologies include finger-reading technology, where users use their fingers to point to corresponding locations on some reading materials. Finger-reading technology can read aloud the text content contained in the location pointed by the user's finger, or prompt relevant information. Especially in some children's learning scenarios, finger-reading technology can help users with reading assistance, help users understand, and improve reading efficiency.
[0003] Summary of the Invention
[0004] The technical problem to be solved by the present disclosure is to provide a character recognition method and device, which can improve the accuracy of fingertip character positioning and recognition.
[0005] To solve the above technical problems, the embodiments of the present disclosure provide the following technical solutions:
[0006] In one aspect, a character recognition method is provided, comprising:
[0007] A screenshot acquisition step is to acquire a screenshot of the area where the user's fingertip is located;
[0008] an area text box detection step, performing text detection on the acquired screenshot to obtain an area text box, wherein the area text box includes a row text box and / or a column text box;
[0009] A matching step of obtaining a target area text box that matches the user's fingertip point;
[0010] The detection and recognition step performs single-word detection on the target area text box, and determines the target character pointed to by the user's fingertip according to the single-word detection result or according to the single-word detection result and character positioning information.
[0011] In some embodiments, the screenshot acquisition step includes:
[0012] Get an image of the user's finger;
[0013] determining the fingertip point and the finger joint point of the user's finger from the image;
[0014] Calculating a fingertip direction and a fingertip angle according to the fingertip point and the finger joint point of the user's finger, wherein the fingertip angle is the angle between the user's fingertip direction and the opposite direction of the x-axis of the image;
[0015] A screenshot is taken from the image according to the fingertip point angle and the fingertip point location.
[0016] In some embodiments, the relationship between the screenshot, the angle of the fingertip point, and the position of the fingertip point satisfies any of the following:
[0017] When the angle of the fingertip point is greater than or equal to 0° and less than 75°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is two-thirds of the screenshot size;
[0018] When the angle of the fingertip point is greater than or equal to 75° and less than 100°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-half of the screenshot size;
[0019] When the angle of the fingertip point is greater than or equal to 100° and less than 180°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-third of the screenshot size.
[0020] In some embodiments, after the screenshot acquisition step, the method further includes:
[0021] The step of performing trapezoidal correction on the screenshot.
[0022] In some embodiments, the method further comprises:
[0023] When the number of detected area text boxes is 0, perform the following steps:
[0024] Step a: After enlarging the screenshot size, re-execute the screenshot acquisition step;
[0025] Step b: determine whether the number of detected regional text boxes is 0. If so, proceed to step c; if not, do not execute the screenshot acquisition step;
[0026] Step c: Determine whether the number of times the screenshot size has been enlarged reaches a preset number. If not, go to step a. If yes, do not execute the screenshot acquisition step.
[0027] In some embodiments, the area text box is a line text box, and the matching step includes:
[0028] When the user's fingertip point falls into the line text box, determining a target line text box that matches the user's fingertip point based on vertical distances between the user's fingertip point and four edges of the line text box;
[0029] When the user's fingertip point does not fall within the line text box, a target line text box matching the user's fingertip point is determined based on the line text box intersecting with the user's fingertip direction and the vertical distance between the user's fingertip point and the lower edge of the line text box.
[0030] In some embodiments, determining the target line text box that matches the user's fingertip point based on the vertical distance between the user's fingertip point and four edges of the line text box includes:
[0031] When the line text box is a horizontal text box, if the vertical distance between the user's fingertip point and the four edges of the line text box satisfies: db<(dt+db) / 3 and min(dr,dl)>box_h / 4, then it is determined that the line text box matches the user's fingertip point, wherein db is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl is the vertical distance between the user's fingertip point and the left edge of the line text box, dr is the vertical distance between the user's fingertip point and the right edge of the line text box, and box_h is the height of the line text box;
[0032] When the line text box is a curved text box, if the vertical distance between the user's fingertip point and the four edges of the line text box satisfies: db1<(dt1+db1)*2 / 3 and min(dr1,dl1)>box_h 1 / 4, and abs(c_db)<(dt1+db1) / 3, then the line text box is judged to match the user's fingertip point, wherein c_db is the minimum distance from the user's fingertip point to the text polygon outline of the curved text box, db1 is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt1 is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl1 is the vertical distance between the user's fingertip point and the left edge of the line text box, and dr1 is the vertical distance between the user's fingertip point and the right edge of the line text box, and the curved text box satisfies: the ratio of the area of the line text box to the area of the text polygon outline is greater than a preset ratio.
[0033] In some embodiments, the method further comprises:
[0034] When the height of the line text box is greater than a preset first threshold and the vertical distance between the user's fingertip point and the lower edge of the line text box is less than the product of the height of the line text box and a preset first ratio, it is determined that the user's fingertip point falls within the line text box.
[0035] In some embodiments, determining a target row text box that matches the user's fingertip point based on the row text box that intersects the direction of the user's fingertip and the vertical distance between the user's fingertip point and the lower edge of the row text box includes:
[0036] If there is a row text box that intersects with the direction of the user's fingertip, and the vertical distance between the user's fingertip point and the lower edge of the row text box is less than a preset second threshold, the row text box is determined to be the target row text box that matches the user's fingertip point, and the second threshold is determined by the average height of the row text boxes in the screenshot.
[0037] In some embodiments, if there is no row text box that intersects the direction of the user's fingertip, the method further includes:
[0038] Enlarge the size of all text boxes in the screenshot;
[0039] Determine whether there is a row text box that intersects with the direction of the user's fingertip. If not, determine among the row text boxes above the user's fingertip point that the row text box whose vertical distance between the lower edge and the user's fingertip point is less than the second threshold is the target row text box that matches the user's fingertip point.
[0040] In some embodiments, the area text box is a line text box, and after the matching step, the method further includes:
[0041] A judgment step, judging whether a preset image expansion condition is met;
[0042] an image enlargement step, when the image enlargement condition is met, enlarging the size of the screenshot and re-performing the screenshot acquisition step;
[0043] The expansion conditions include any of the following:
[0044] When no target text box matching the user's fingertip point is obtained, the ratio of the maximum height of all text boxes in the screenshot to the screenshot size is greater than a preset second ratio;
[0045] When a target row text box matching the user's fingertip point is obtained, the ratio of the minimum vertical distance between the target row text box and the edge of the screenshot to the screenshot size is less than a preset third ratio, and the ratio of the height of the target row text box to the screenshot size is greater than a preset second ratio.
[0046] In some embodiments, the area text box is a line text box, and the detection and identification step includes:
[0047] Acquire a text line image to be identified corresponding to the target text line box from the screenshot according to the position information of the target text line box;
[0048] Obtaining a character positioning point from the to-be-recognized line text image according to the direction of the user's fingertip;
[0049] Performing single-word detection on the line text image to be recognized to obtain a plurality of single-word detection frames, and removing single-word detection frames that meet preset requirements from the plurality of single-word detection frames, wherein the preset requirements are: a width of the single-word detection frame is less than a product of an average width of the single-word detection frame and a preset magnification, and a minimum distance between the single-word detection frame and an edge of the line text image to be recognized is less than a preset third threshold;
[0050] Cropping the to-be-recognized line text image according to the removed single-word detection frame, and performing text recognition on the cropped to-be-recognized line text image to obtain a recognized text including a plurality of characters;
[0051] If the number of characters included in the recognized text is equal to the number of the single-character detection frames, determining a target single-character detection frame corresponding to the character positioning point, and determining a target character among the multiple characters according to the sequence number of the target single-character detection frame in the multiple single-character detection frames;
[0052] If the number of characters included in the recognized text is not equal to the number of the single-word detection frames, the connected text classification CTC information in the convolutional recurrent neural network CRNN recognition network is used to estimate the character position, and the target character corresponding to the character positioning point is output.
[0053] In some embodiments, obtaining a character positioning point from the to-be-recognized line text image according to the direction of the user's fingertip includes:
[0054] When the target line text box is a horizontal text box, when the acute angle formed by the fingertip direction and the lower edge of the target line text box is greater than 45°, determining the intersection of a first straight line and the fingertip direction, where the first straight line intersects the target line text box, is parallel to the lower edge, and the distance between the first straight line and the lower edge is one-third of the height of the target line text box; when the acute angle formed by the fingertip direction and the lower edge of the target line text box is less than or equal to 45°, determining the intersection of the lower edge and the fingertip direction;
[0055] When the target line text box is a curved text box, when the fingertip point is located within the character polygon outline of the target line text box, the fingertip point is determined as the intersection point; when the fingertip point is located within the target line text box and outside the character polygon outline, the intersection point of a second straight line and the fingertip direction is determined, where the second straight line is parallel to the lower edge of the target line text box and the distance between the second straight line and the lower edge and the upper edge of the target line text box is half of the height of the target line text box; when the fingertip point is located outside the target line text box, the intersection point of the second straight line and the fingertip direction is determined;
[0056] The coordinates of the intersection are mapped to the line text image to be recognized to obtain the character positioning point.
[0057] In some embodiments, after obtaining multiple single-word detection frames, the method further includes:
[0058] When the distance between two adjacent single-word detection frames is greater than a preset fourth threshold, a virtual single-word detection frame is added between the two adjacent single-word detection frames, and the fourth threshold is determined by the average width of the single-word detection frames.
[0059] In some embodiments, after obtaining multiple single-word detection frames, the method further includes:
[0060] When the target line text box is a curved text box, if a preset condition is met, retain S single-word detection boxes closest to the character positioning point, and shrink the line text image to be recognized according to the retained S single-word detection boxes, where S is a positive integer;
[0061] The preset conditions include:
[0062] The ratio of the aspect ratio of the to-be-recognized line text image to the number of the single-word detection frames is between 0.9 and 1.8, and the number of the single-word detection frames is greater than S.
[0063] In some embodiments, the method further comprises:
[0064] If the target character is a punctuation mark or a space, the target character is replaced with the character that is the nearest non-punctuation mark or space to the target character.
[0065] In some embodiments, the method further comprises:
[0066] If the target character is English, the recognized text is searched forward and backward starting from the target character until a punctuation mark or space is encountered, to obtain an alternative English character string, segment the alternative English character string, and replace the target character with the segmentation result.
[0067] An embodiment of the present disclosure further provides a character recognition device, comprising:
[0068] A screenshot acquisition module is used to obtain a screenshot of the area where the user's fingertip is located;
[0069] A regional text box detection module is used to perform text detection on the obtained screenshot to obtain a regional text box, wherein the regional text box includes a row text box and / or a column text box;
[0070] A matching module, configured to obtain a target area text box that matches the user's fingertip point;
[0071] The detection and recognition module is used to perform single-word detection on the target area text box, and determine the target character pointed to by the user's fingertip according to the single-word detection result or according to the single-word detection result and character positioning information.
[0072] An embodiment of the present disclosure further provides a character recognition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the program, the character recognition method described above is implemented.
[0073] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned character recognition method when executed by a processor.
[0074] The embodiments of the present disclosure have the following beneficial effects:
[0075] In the above scheme, after performing text detection on the screenshot of the area where the user's fingertip is located, the regional text box is obtained, and then single-word detection is performed on the regional text box. The target character pointed by the user's fingertip is determined based on the single-word detection result. This embodiment adopts a scheme that combines single-word detection with text recognition results, which can improve the accuracy of fingertip character positioning and recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] FIG1 is a flow chart of a character recognition method according to an embodiment of the present disclosure;
[0077] FIG2 is a schematic diagram of a process of determining a target row text box that matches a user's fingertip point when the user's fingertip point falls within a row text box according to an embodiment of the present disclosure;
[0078] FIG3 is a schematic diagram of a process of determining a target row text box matching a user's fingertip point when the user's fingertip point falls outside a row text box according to an embodiment of the present disclosure;
[0079] FIG4 is a schematic diagram showing the effect of enlarging a screenshot according to an embodiment of the present disclosure;
[0080] FIG5 is a schematic diagram of a process for determining a target character pointed to by a user's fingertip according to an embodiment of the present disclosure;
[0081] FIG6 is a schematic diagram of multiple original single-word detection frames according to an embodiment of the present disclosure;
[0082] FIG7 is a schematic diagram of an embodiment of the present disclosure after removing the truncated single-word detection frame;
[0083] FIG8 is a schematic diagram of multiple single-word detection frames before shrinkage according to an embodiment of the present disclosure;
[0084] FIG9 is a schematic diagram of multiple single-word detection frames after shrinkage according to an embodiment of the present disclosure;
[0085] FIG10 is a schematic diagram showing an embodiment of the present disclosure in which the number of characters included in a recognized text is equal to the number of single-word detection frames;
[0086] FIG11 is a schematic diagram showing an embodiment of the present disclosure in which the number of characters included in a recognized text is not equal to the number of single-word detection frames;
[0087] FIG12 is a schematic structural diagram of a character recognition device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0088] In order to make the technical problems, technical solutions and advantages to be solved by the embodiments of the present disclosure more clear, they will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0089] In reading assistance technology, users point their finger at a specific location in a reading material. The device's camera then takes a picture of the area where the user's finger is located. The device analyzes the image to identify the character pointed at, then reads the character aloud or provides relevant information. However, due to factors such as curved text and incomplete capture of large text, the accuracy of fingertip character location and recognition can be low.
[0090] The embodiments of the present disclosure provide a character recognition method and device, which can improve the accuracy of fingertip character positioning and recognition.
[0091] An embodiment of the present disclosure provides a character recognition method, as shown in FIG1 , comprising:
[0092] Screenshot acquisition step 101, acquiring a screenshot of the area where the user's fingertip is located;
[0093] The purpose of finger reading is to determine the single Chinese character, single English word or other type of character pointed to by the user's fingertip. Therefore, after taking an image including the reading material and the user's finger, there is no need to recognize the entire image. A screenshot can be taken from the captured image, and text detection and recognition can be performed on the screenshot. This can reduce the amount of calculation and increase the calculation speed.
[0094] When taking a screenshot for the first time, you can use the preset size img_size. That is, the length of the screenshot is img_size, and the width is also img_size.
[0095] Specifically, the fingertip point p1 and the fingertip node p2 of the user's finger are determined from the captured image; the fingertip direction vector P and the fingertip point angle angle are calculated based on the fingertip point p1 and the fingertip node p2 of the user's finger, wherein a plane coordinate system is established in the image, the x-axis direction is the row direction of the text in the reading material, the y-axis is perpendicular to the x-axis, and the fingertip point angle is the angle between the user's fingertip direction and the x-axis; a screenshot is taken from the image based on the fingertip point angle and the position of the fingertip point.
[0096] Specifically, the relationship between the screenshot, the angle of the fingertip point, and the position of the fingertip point satisfies any of the following:
[0097] When the angle of the fingertip point is greater than or equal to 0° and less than 75°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is two-thirds of the screenshot size;
[0098] When the angle of the fingertip point is greater than or equal to 75° and less than 100°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-half of the screenshot size;
[0099] When the angle of the fingertip point is greater than or equal to 100° and less than 180°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-third of the screenshot size.
[0100] That is, if angle∈[0,75), take the fingertip point p1 as the starting point, extend two-thirds of the screenshot size img_size to the left, one-third of the screenshot size img_size to the right, one-quarter of the screenshot size img_size to the bottom, and three-quarters of the screenshot size img_size to the top to take a screenshot; if angle∈[100,180), take the fingertip point p1 as the starting point, extend one-third of the screenshot size img_size to the left, and one-third of the screenshot size img_size to the right If angle∈[75,100), take the screenshot starting from the fingertip point p1 and extend it to the left by half the size of img_size, to the right by half the size of img_size, to the bottom by one-quarter the size of img_size, and to the top by three-quarters the size of img_size. This ensures that the captured screenshot includes the character pointed to by the user's fingertip.
[0101] Generally, the camera uses a top-down perspective to shoot, and the captured image has trapezoidal distortion. Compared with vertically shot images, text detection is more difficult. Therefore, after obtaining the screenshot, the method also includes: a trapezoidal correction step, which performs trapezoidal correction on the screenshot. This can improve the accuracy of character recognition based on the screenshot.
[0102] Screenshots are generally rectangular. When performing trapezoidal correction, the perspective transformation matrix H of the entire image can be estimated based on the camera intrinsic parameters and camera angle. Combined with the four-point coordinates of the screenshot rectangle, the perspective transformation matrix M of the corresponding screenshot is calculated, and the screenshot is then subjected to trapezoidal correction using the perspective transformation matrix M. Camera intrinsic parameters include focal length, lens blur radius, and image quality description parameters. Focal length represents the relationship between image size and focal length; lens blur radius refers to the distance from the optical center to the edge of the image; and image quality description parameters include indicators such as brightness range and noise level. The camera angle is the angle between the camera's shooting direction and the z-axis, which is perpendicular to the plane containing the x- and y-axes.
[0103] Regional text box detection step 102, performing text detection on the obtained screenshot to obtain regional text boxes, wherein the regional text boxes include row text boxes and / or column text boxes;
[0104] The characters in the reading material can be arranged in rows or columns. When the characters are arranged in rows, it is necessary to perform row text detection on the captured screenshot to obtain a row text box; when the characters are arranged in columns, it is necessary to perform column text detection on the captured screenshot to obtain a column text box.
[0105] In this embodiment, a DBNet model can be used to perform text detection on the captured screenshot to obtain a regional text box. If no regional text box is detected, it may be because the screenshot size is too small and the text exceeds the screenshot range, there may be no text in the screenshot, or the text in the screenshot is too small to be detected. The screenshot can be enlarged and retaken, and text detection can be performed again. If the regional text box is still not detected, the screenshot size can be further enlarged until the preset expansion times EXPAND_TIMES are reached. If the regional text box is still not detected, it is assumed that the user's fingertip is not pointing to any characters.
[0106] Specifically, when the number of detected area text boxes is 0, perform the following steps:
[0107] Step a: After enlarging the screenshot size, re-execute the screenshot acquisition step;
[0108] The screenshot size can be expanded according to the preset expansion ratio EXPAND_RATIO, that is, the screenshot size of each screenshot is EXPAND_RATIO times the screenshot size of the previous screenshot. EXPAND_RATIO is greater than 1 and can be 1.5, 2, 2.5 or 3.
[0109] Step b: determine whether the number of detected regional text boxes is 0. If so, proceed to step c; if not, do not execute the screenshot acquisition step;
[0110] Step c: Determine whether the number of times the screenshot size has been enlarged reaches a preset number. If not, go to step a. If yes, do not execute the screenshot acquisition step.
[0111] Matching step 103, obtaining a target area text box that matches the user's fingertip point;
[0112] Taking the regional text box as an example, we can first determine whether the user's fingertip point is located in a row text box. If the user's fingertip point p1 falls within multiple row text boxes det_box, we determine a row text box det_box from the multiple row text boxes det_box whose lower edge (i.e., bottom edge) is closest to the user's fingertip point p1 and use this row text box det_box as the row text box where the user's fingertip point p1 is located.
[0113] When the user's fingertip point falls within the line text box, as shown in Figure 2, the target line text box that matches the user's fingertip point can be determined based on the vertical distance between the user's fingertip point and the four edges of the line text box. The line text box where the user's fingertip point p1 is located is rec_box, and the height of rec_box needs to be greater than a preset first threshold. If the height of rec_box is too small, such as less than 8 pixels, the rec_box can be ignored. If character recognition is performed on the rec_box, the recognized characters will be blurred due to their small size, and the recognition accuracy cannot be guaranteed.
[0114] Specifically, when the height box_h of the line text box is greater than a preset first threshold value BIG_CHAR_THRE and the vertical distance between the user's fingertip point and the bottom edge of the line text box is less than the product of the height box_h of the line text box and a preset first ratio INBOX_RATIO, it can be determined that the user's fingertip point falls within the line text box. That is, when the line text box is a large line text box and the user's fingertip point p1 is close to the bottom edge of rec_box, it is determined that the user's fingertip point falls within the line text box. The value of the first threshold value BIG_CHAR_THRE can be set as needed, for example, to 32 pixels, and the first ratio INBOX_RATIO can be set as needed, for example, to 0.3.
[0115] After determining that the user's fingertip point falls within the line text box, the target line text box that matches the user's fingertip point can be further determined based on the vertical distance between the user's fingertip point and the four edges of the line text box. When the line text box is a horizontal text box, if the vertical distance between the user's fingertip point and the four edges of the line text box meets the first condition: db<(dt+db) / 3 and min(dr,dl)>box_h / 4, then it is determined that the line text box matches the user's fingertip point, wherein db is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl is the vertical distance between the user's fingertip point and the left edge of the line text box, dr is the vertical distance between the user's fingertip point and the right edge of the line text box, and box_h is the height of the line text box, that is, the user's fingertip point needs to be deviated to the bottom of the line text box and maintain a certain distance from the left and right edges of the line text box. If this is met, it is confirmed that the user's fingertip point points to the current rec_box.
[0116] When the line text box is a curved text box, in order to avoid line text box matching failure and line skipping recognition, it is necessary to change the judgment condition. A curved text box means that the text in the line text box is not arranged in a straight line due to the curvature of the text. First, based on the line text box, the polygonal outline of the text in the line text box can be extracted. When the ratio of the area of the line text box (Area (rec_box)) to the area of the polygonal outline of the text (Area (rec_poly)) is greater than the preset ratio CURVED_RATIO_THRE, the line text box is judged to be a curved text box. After determining that the line text box is a curved text box, if the vertical distance between the user's fingertip point and the four edges of the line text box meets the second condition: db1<(dt1+db1)*2 / 3 and min(dr1,dl1)>box_h 1 / 4, and abs(c_db)<(dt1+db1) / 3, then the line text box is determined to match the user's fingertip point, where c_db is the minimum distance from the user's fingertip point to the text polygon outline of the curved text box, db1 is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt1 is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl1 is the vertical distance between the user's fingertip point and the left edge of the line text box, and dr1 is the vertical distance between the user's fingertip point and the right edge of the line text box.
[0117] When the user's fingertip point does not fall within the row text box, as shown in FIG3 , the target row text box that matches the user's fingertip point can be determined based on the row text box that intersects with the user's fingertip direction and the vertical distance between the user's fingertip point and the lower edge of the row text box. Specifically, if there is a row text box rec_box that intersects with the user's fingertip direction, and the vertical distance between the user's fingertip point and the lower edge of the row text box rec_box is less than the preset second threshold p2l_thre, the row text box is determined to be the target row text box that matches the user's fingertip point, and the second threshold p2l_thre is determined by the mean height of the row text boxes in the screenshot. Specifically, the second threshold p2l_thre = min(max(mean_h,20),30)*OUT_BOX_THRE, OUT_BOX_THRE is the distance multiplier threshold for determining the direction to rec_box, and the value can be 1.5. If there is no row text box that intersects with the direction of the user's fingertip, the size of all row text boxes in the screenshot can be expanded, for example, all det_boxes can be expanded by half a character height (box_h) on the left and right, and then it is determined again whether there is a row text box that intersects with the direction of the user's fingertip. If not, the row text box above the user's fingertip point is determined to be the target row text box that matches the user's fingertip point, with the vertical distance between the bottom edge and the user's fingertip point being less than the second threshold p2l_thre. If there are multiple row text boxes with the vertical distance between the bottom edge and the user's fingertip point being less than the second threshold p2l_thre, the row text box closest to the user's fingertip point is selected.
[0118] In this embodiment, text recognition is performed on the screenshot crop_img. For text with a large pixel area, truncation is likely to occur, or text may be missed or the line text box may be incomplete. Therefore, it is necessary to determine whether the text of the current reading material is large characters based on the line text box height, that is, to determine whether the preset image expansion conditions are met. If so, the image expansion is performed again. Specifically, after the matching step 103 and before the detection and recognition step 104, the following steps may also be included:
[0119] A judgment step, judging whether a preset image expansion condition is met;
[0120] an image enlargement step, when the image enlargement condition is met, enlarging the size of the screenshot and re-performing the screenshot acquisition step;
[0121] The expansion conditions include any of the following:
[0122] When no target text box matching the user's fingertip point is obtained, the ratio of the maximum height of all text boxes in the screenshot to the screenshot size is greater than a preset second ratio;
[0123] When a target row text box matching the user's fingertip point is obtained, the ratio of the minimum vertical distance between the target row text box and the edge of the screenshot to the screenshot size is less than a preset third ratio, and the ratio of the height of the target row text box to the screenshot size is greater than a preset second ratio.
[0124] If step 103 does not find the rec_box that matches the user's fingertip point, a large character judgment is performed, and the maximum height max_box_h of all text boxes in the screenshot is calculated. If the ratio of max_box_h to the screenshot height crop_img_h is greater than the preset second ratio BIG_CHAR_RATIO_THRE, it is considered that large characters exist, and the expansion ratio is set to min(max(EXPAND_RATIO*max_box_h / BIG_CHAR_THRE,2),MAX_EXPAND_RATIO), where BIG_CHAR_THRE is the pixel threshold for large character judgment, the initial value of EXPAND_RATIO is 1, and MAX_EXPAND_RATIO is the limited maximum expansion ratio.
[0125] If step 103 finds a rec_box that matches the user's fingertip point, in order to avoid unnecessary image expansion, large character judgment is only performed for the case where the rec_box is close to the edge of the screenshot. If the ratio of the minimum vertical distance between the target row text box rec_box and the edge of the screenshot to the screenshot size is less than a preset third ratio, it is determined that the target row text box rec_box is located at the edge of the screenshot. Whether the current rec_box is a large character text box is determined based on whether the ratio of the height box_h of the rec_box to the screenshot height crop_img_h is greater than the preset second ratio BIG_CHAR_RATIO_THRE.
[0126] After the screenshot area is enlarged, steps 101 to 103 are executed again. FIG4 is a schematic diagram showing the effect of enlarging the screenshot.
[0127] When the area text box is a column text box, it can be first determined whether the user's fingertip point is located in a column text box. If the user's fingertip point p1 falls within multiple column text boxes det_box, a column text box det_box is determined from the multiple column text boxes det_box, and the left edge or right edge of the column text box det_box is closest to the user's fingertip point p1. The column text box det_box is used as the column text box where the user's fingertip point p1 is located.
[0128] When the user's fingertip point falls within the column text box, the target column text box that matches the user's fingertip point can be determined based on the vertical distance between the user's fingertip point and the four edges of the column text box. The column text box where the user's fingertip point p1 is located is called rec_box. The width of rec_box needs to be greater than a preset threshold. Because if the width of rec_box is too small, such as less than 8 pixels, the rec_box can be ignored. If character recognition is performed on the rec_box, the recognized characters will be blurred due to the small size, and the recognition accuracy cannot be guaranteed.
[0129] Specifically, when the width of the column text box is greater than a preset threshold BIG_CHAR_THRE and the vertical distance between the user's fingertip point and the left edge or right edge of the column text box is less than the product of the width of the column text box and the preset first ratio INBOX_RATIO, it can be determined that the user's fingertip point falls within the column text box. That is, when the column text box is a large column text box and the user's fingertip point p1 is close to the left edge or right edge of the column text box, it is determined that the user's fingertip point falls within the column text box. The value of the preset threshold BIG_CHAR_THRE can be set as needed, for example, to 32 pixels, and the first ratio INBOX_RATIO can be set as needed, for example, to 0.3.
[0130] After determining that the user's fingertip point falls within the column text box, the target column text box matching the user's fingertip point may be further determined based on the vertical distance between the user's fingertip point and the four edges of the column text box.
[0131] In the detection and recognition step 104 , single-word detection is performed on the target area text box, and the target character pointed to by the user's fingertip is determined based on the single-word detection result or based on the single-word detection result and character positioning information.
[0132] Taking the area text box as a row text box as an example, as shown in FIG5 , the detection and recognition step 104 includes the following steps:
[0133] (1) First, the image of the line of text to be identified corresponding to the target line of text box can be obtained from the screenshot according to the position information of the target line of text box, and the coordinates of the rec_box matched with the user's fingertip point determined in step 103 are mapped back to the original image, that is, the screenshot before the trapezoidal correction is performed, and the image of the line of text to be identified, crop_line_img, is captured from the screenshot. Capturing the image from the screenshot before the trapezoidal correction can avoid the blurring of the text image caused by the difference in the trapezoidal correction.
[0134] (2) obtaining a character positioning point from the to-be-recognized line text image according to the direction of the user's fingertip;
[0135] In this embodiment, different methods are used to determine character positioning points for horizontal text boxes and curved text boxes. A curved text box is one in which the characters are not arranged in a straight line due to the curve of the text. Based on the line text box, the polygonal outline of the characters within the line text box can be extracted. When the ratio of the area of the line text box (Area (rec_box)) to the area of the polygonal outline of the characters (Area (rec_poly)) is greater than a preset ratio (CURVED_RATIO_THRE), the line text box is determined to be a curved text box.
[0136] When the target line text box is a horizontal text box, when the acute angle formed by the fingertip direction and the lower edge of the target line text box is greater than 45°, determining the intersection of a first straight line and the fingertip direction, where the first straight line intersects the target line text box, is parallel to the lower edge, and the distance between the first straight line and the lower edge is one-third of the height of the target line text box; when the acute angle formed by the fingertip direction and the lower edge of the target line text box is less than or equal to 45°, determining the intersection of the lower edge and the fingertip direction;
[0137] When the target line text box is a curved text box, when the fingertip point is located within the character polygon outline of the target line text box, the fingertip point is determined as the intersection point; when the fingertip point is located within the target line text box and outside the character polygon outline, the intersection point of a second straight line and the fingertip direction is determined, where the second straight line is parallel to the lower edge of the target line text box and the distance between the second straight line and the lower edge and the upper edge of the target line text box is half of the height of the target line text box; when the fingertip point is located outside the target line text box, the intersection point of the second straight line and the fingertip direction is determined;
[0138] The coordinates of the intersection are mapped to the crop_line_img line text image to be recognized to obtain the character positioning point.
[0139] (3) performing single word detection on the text image to be recognized to obtain multiple single word detection frames;
[0140] Specifically, the yolov5 deep learning network can be used to perform single-word detection on the crop_line_img line text image to be identified. This embodiment performs single-word detection on the basis of line text detection, which is time-saving, reduces the difficulty of detection, and does not need to consider the direction of the text.
[0141] Since punctuation marks are easily missed, after obtaining multiple single-word detection frames, the distance between two adjacent single-word detection frames is analyzed. When the distance between two adjacent single-word detection frames is greater than a preset fourth threshold, a virtual single-word detection frame is added between the two adjacent single-word detection frames. The fourth threshold is determined by the average width of the single-word detection frames, and the fourth threshold can be equal to the product of the average width of the single-word detection frames and a preset multiplier r.
[0142] In this embodiment, a screenshot is taken during the text detection stage, so the first and last characters of the line text are easily truncated, and the recognition of truncated characters is unstable. Depending on the degree of truncation, recognition errors may occur or the single-word detection frame may not be recognized. When the single-word detection frame cannot be recognized, the number of single-word detection frames does not match the number of characters included in the recognized text rec_text, affecting the accuracy of character positioning. Therefore, it is necessary to remove the truncated single-word detection frames from the single-word detection frames, that is, to remove the single-word detection frames that meet the preset requirements from the multiple single-word detection frames. The preset requirements are: the width of the single-word detection frame is less than the product of the average width of the single-word detection frame and the preset magnification, and the minimum distance between the single-word detection frame and the edge of the line text image to be recognized is less than a preset third threshold.
[0143] Specifically, the coordinates of the single-word detection box inv_char_box mapped back to the original image can be used to determine whether the single-word detection box is close to the edge of the crop_line_img line text image to be identified. For the single-word detection box close to the edge of the crop_line_img line text image to be identified, the head and tail truncation judgment is performed, because only the single-word detection box at the edge of the crop_line_img line text image to be identified will be truncate. For the single-word detection frame close to the left edge of the line text image crop_line_img to be identified, that is, at the beginning of the line text, if the width of the first single-word detection frame is less than the product of the average width mean_w of the single-word detection frame and r_c, and the distance between it and the left edge of the line text image crop_line_img to be identified is less than mean_w / 2, then the first single-word detection frame is removed, wherein r_c is a preset magnification, such as 0.8; for the single-word detection frame close to the right edge of the line text image crop_line_img to be identified, that is, at the end of the line text, if the width of the last single-word detection frame is less than mean_w*r_c, and the distance between it and the right edge of the line text image crop_line_img to be identified is less than mean_w / 2, then the last single-word detection frame is removed. As shown in Figure 6, a schematic diagram of the original multiple single-word detection frames in a specific example, Figure 7 is a schematic diagram after removing the truncated single-word detection frames.
[0144] In addition, in the picture book finger reading scenario, due to the natural curvature of the book, curved text is easily generated. To improve the accuracy of curved text recognition, a specific character length can be set based on the single-word detection frame to avoid some curved text. When the text length is reduced, the corresponding curvature is reduced, and the background area is reduced, which can improve the accuracy of character recognition. Specifically, when the target line text frame is a curved text frame, if the preset conditions are met, the S single-word detection frames closest to the character positioning point are retained, and the line text image to be recognized (crop_line_img) is shrunk based on the retained S single-word detection frames to obtain the shrunk line text image (shrink_img), where S is a positive integer;
[0145] The preset conditions include:
[0146] The ratio of the aspect ratio of the to-be-recognized line text image to the number of the single-word detection frames is between 0.9 and 1.8, and the number of the single-word detection frames is greater than S.
[0147] The purpose of determining the aspect ratio of the text line to be recognized is to avoid truncating English words, as English characters differ from Chinese characters, and English words are output. Figure 8 shows a schematic diagram of multiple single-word detection frames before contraction in a specific example, and Figure 9 shows a schematic diagram of multiple single-word detection frames after contraction.
[0148] (4) cropping the line text image to be recognized according to the removed single-word detection frame, and performing text recognition on the cropped line text image to be recognized to obtain a recognized text including multiple characters;
[0149] Specifically, a CRNN model can be used to obtain the recognized text rec_text. When the character recognition device of this embodiment is deployed in the cloud, a large model can be used for recognition. When the character recognition device of this embodiment is deployed in the terminal, a lightweight model can be used for recognition.
[0150] (5) If the number of characters included in the recognized text is equal to the number of the single-word detection frames, determine the target single-word detection frame corresponding to the character positioning point, and determine the target character among the multiple characters according to the sequence number of the target single-word detection frame in the multiple single-word detection frames; if the number of characters included in the recognized text is not equal to the number of the single-word detection frames, estimate the character position using the connected text classification CTC information of the recognition network, and output the target character corresponding to the character positioning point.
[0151] Figure 10 shows a schematic diagram of a specific example in which the number of characters included in the recognized text is equal to the number of the single-word detection frames. When the number of characters included in the recognized text is equal to the number of the single-word detection frames, the multiple single-word detection frames are sorted according to the x-axis coordinates of the single-word detection frames, and the distance between the center point of each single-word detection frame and the character positioning point can be calculated. The single-word detection frame with the closest center point to the character positioning point is determined as the target single-word detection frame, and the target character among the multiple characters is determined according to the sequence number of the target single-word detection frame in the multiple single-word detection frames. For example, if the sequence number of the target single-word detection frame in the multiple single-word detection frames is 5, then the fifth character among the multiple characters is determined to be the target character; if the sequence number of the target single-word detection frame in the multiple single-word detection frames is 3, then the third character among the multiple characters is determined to be the target character, and so on. This embodiment adopts a solution of matching the single-word detection results with the line text recognition results, which can improve the accuracy of fingertip character positioning and recognition.
[0152] FIG11 is a schematic diagram showing a specific example in which the number of characters included in the recognized text is not equal to the number of single-word detection frames due to missed detection of punctuation marks. In this embodiment, the character positioning information can specifically be CTC information, which can be used to determine the positioning information of the character. When the number of characters included in the recognized text is not equal to the number of the single-word detection frames, the character position can be estimated based on the CTC information in the text recognition result. Specifically, the character positioning point is converted into an index cross_idx, and the index cross_idx is determined by the x-axis coordinate C_X of the character positioning point in the screenshot and the subsampling number stride of the input image width of the recognition network. Specifically, cross_idx=C_X / / stride, / / represents rounding, and the subsampling number stride of the input image width is the input image slice width corresponding to each bit in the vector Y obtained by softmax calculation of the CTC information output probability map.
[0153] According to the CTC decoding rules, 'blank' is a placeholder, and the index is decoded. If Y[cross_idx] is decoded as 'blank', the difference dist = C_X-cross_idx*stride is calculated. If dist is less than 0, starting from the placeholder blank, the nearest non-blank character is searched in the recognized text in the order of searching one character to the left and then searching one character to the right to obtain the target character; if dist is greater than or equal to 0, the nearest non-blank character is searched in the recognized text in the order of searching one character to the right and then searching one character to the left to obtain the target character.
[0154] In this embodiment, when the single-word detection result does not match the line text recognition result, the CTC information is used to estimate the target character position, thereby ensuring the accurate output of the target character.
[0155] (6) If the target character is a punctuation mark or a space, since punctuation marks or spaces are meaningless characters, the target character is replaced with the character that is the nearest non-punctuation mark or space to the target character.
[0156] (7) If the target character is in English, the recognized text is searched forward and backward starting from the target character until a punctuation mark or a space is encountered, and an alternative English character string is obtained. The alternative English character string is segmented, and the target character is replaced with the segmentation result.
[0157] In addition, the execution order of the above steps can also be (1)(2)(4)(3)(5)(6)(7), or (1)(4)(2)(3)(5)(6)(7).
[0158] When the area text box is a column text box, the detection and identification step 104 includes the following steps:
[0159] (1) First, the to-be-recognized column text image corresponding to the target column text box can be obtained from the screenshot according to the position information of the target column text box, and the coordinates of the rec_box matched with the user's fingertip point determined in step 103 are mapped back to the original image, that is, the screenshot before the trapezoidal correction is performed, and the to-be-recognized column text image crop_line_img is captured from the screenshot. Capturing the image from the screenshot before the trapezoidal correction can avoid the text image blurring caused by the difference in the trapezoidal correction.
[0160] (2) obtaining a character positioning point from the to-be-recognized text image according to the direction of the user's fingertip;
[0161] (3) performing single word detection on the text image to be identified to obtain multiple single word detection frames;
[0162] Specifically, the yolov5 deep learning network can be used to perform single-word detection on the column text image crop_line_img to be identified. This embodiment performs single-word detection on the basis of column text detection, which is time-saving, reduces the difficulty of detection, and does not need to consider the direction of the text.
[0163] Since punctuation marks are easily missed, after obtaining multiple single-word detection frames, the distance between two adjacent single-word detection frames is analyzed. When the distance between two adjacent single-word detection frames is greater than a preset fourth threshold, a virtual single-word detection frame is added between the two adjacent single-word detection frames. The fourth threshold is determined by the average height of the single-word detection frames, and the fourth threshold can be equal to the product of the average height of the single-word detection frames and a preset magnification r.
[0164] In this embodiment, a screenshot is taken during the text detection stage, so the first and last characters of the column text are easily truncated, and the recognition of truncated characters is unstable. Depending on the degree of truncation, recognition errors may occur or the single-word detection frame may not be recognized. When the single-word detection frame cannot be recognized, the number of single-word detection frames does not match the number of characters included in the recognized text rec_text, affecting the accuracy of character positioning. Therefore, it is necessary to remove the truncated single-word detection frames from the single-word detection frames, that is, to remove the single-word detection frames that meet the preset requirements from the multiple single-word detection frames. The preset requirements are: the height of the single-word detection frame is less than the product of the average height of the single-word detection frame and the preset magnification, and the minimum distance between the single-word detection frame and the edge of the column text image to be recognized is less than a preset third threshold.
[0165] Specifically, the coordinates of the single-word detection box inv_char_box mapped back to the original image can be used to determine whether the single-word detection box is close to the edge of the to-be-recognized column text image crop_line_img. For the single-word detection box close to the edge of the to-be-recognized column text image crop_line_img, the head and tail truncation judgment is performed, because only the single-word detection box at the edge of the to-be-recognized column text image crop_line_img will be truncate. For the single-word detection box close to the upper edge of the column text image crop_line_img to be identified, that is, the single-word detection box located at the beginning of the column text, if the height of the first single-word detection box is less than the product of the average height of the single-word detection box and r_c, and the distance between it and the upper edge of the column text image crop_line_img to be identified is less than half of the average height, then the first single-word detection box is removed, where r_c is a preset magnification, such as 0.8; for the single-word detection box close to the lower edge of the column text image crop_line_img to be identified, that is, the single-word detection box located at the end of the column text, if the height of the last single-word detection box is less than the average height * r_c, and the distance between it and the lower edge of the column text image crop_line_img to be identified is less than the average height / 2, then the last single-word detection box is removed.
[0166] (4) cropping the to-be-recognized text image according to the removed single-word detection frame, and performing text recognition on the cropped to-be-recognized text image to obtain a recognized text including a plurality of characters;
[0167] Specifically, a CRNN model can be used to obtain the recognized text rec_text. When the character recognition device of this embodiment is deployed in the cloud, a large model can be used for recognition. When the character recognition device of this embodiment is deployed in the terminal, a lightweight model can be used for recognition.
[0168] (5) If the number of characters included in the recognized text is equal to the number of the single-word detection frames, determine the target single-word detection frame corresponding to the character positioning point, and determine the target character among the multiple characters according to the sequence number of the target single-word detection frame in the multiple single-word detection frames; if the number of characters included in the recognized text is not equal to the number of the single-word detection frames, estimate the character position using the connected text classification CTC information of the recognition network, and output the target character corresponding to the character positioning point.
[0169] When the number of characters included in the recognized text is equal to the number of the single-word detection frames, the multiple single-word detection frames are sorted according to the y-axis coordinates of the single-word detection frames, and the distance between the center point of each single-word detection frame and the character positioning point can be calculated. The single-word detection frame whose center point is closest to the character positioning point is determined to be the target single-word detection frame, and the target character among the multiple characters is determined according to the sequence number of the target single-word detection frame in the multiple single-word detection frames. For example, if the sequence number of the target single-word detection frame in the multiple single-word detection frames is 5, then the 5th character among the multiple characters is determined to be the target character; if the sequence number of the target single-word detection frame in the multiple single-word detection frames is 3, then the 3rd character among the multiple characters is determined to be the target character, and so on. This embodiment adopts a solution of matching the single-word detection results with the column text recognition results, which can improve the accuracy of fingertip character positioning and recognition.
[0170] (6) If the target character is a punctuation mark or a space, since punctuation marks or spaces are meaningless characters, the target character is replaced with the character that is the nearest non-punctuation mark or space to the target character.
[0171] The technical solution of this embodiment can be applied to finger-pointing picture book reading scenarios, and is suitable for devices such as picture book robots and reading tablets. The camera of the character recognition device can take a picture of the area where the user's finger is located, and through image analysis, the target character pointed at by the user's fingertip can be read aloud or related information can be prompted. For example, if the target character pointed at by the user's fingertip is Chinese, the Chinese meaning can be explained or the Chinese word can be formed. For another example, if the target character pointed at by the user's fingertip is English, the Chinese equivalent of the English word can be prompted.
[0172] In this embodiment, after performing text detection on the screenshot of the area where the user's fingertip is located, a regional text box is obtained, and then single-word detection is performed on the regional text box. The target character pointed by the user's fingertip is determined based on the single-word detection result. This embodiment adopts a solution that combines single-word detection with text recognition results, which can improve the accuracy of fingertip character positioning and recognition.
[0173] The embodiment of the present disclosure further provides a character recognition device, as shown in FIG12 , comprising:
[0174] A screenshot acquisition module 21 is used to acquire a screenshot of the area where the user's fingertip is located;
[0175] A region text box detection module 22 is configured to perform text detection on the acquired screenshot to obtain a region text box, wherein the region text box includes a row text box and / or a column text box;
[0176] A matching module 23 is used to obtain a target area text box that matches the user's fingertip point;
[0177] The detection and recognition module 24 is configured to perform single-word detection on the target area text box, and determine the target character pointed to by the user's fingertip according to the single-word detection result or according to the single-word detection result and character positioning information.
[0178] In this embodiment, after performing text detection on the screenshot of the area where the user's fingertip is located, a regional text box is obtained, and then single-word detection is performed on the regional text box. The target character pointed by the user's fingertip is determined based on the single-word detection result. This embodiment adopts a solution that combines single-word detection with text recognition results, which can improve the accuracy of fingertip character positioning and recognition.
[0179] In some embodiments, the screenshot acquisition module 21 is specifically used to acquire an image of the user's finger; determine the fingertip point and finger node point of the user's finger from the image; calculate the fingertip direction and fingertip point angle based on the fingertip point and finger node point of the user's finger, the fingertip point angle being the angle between the user's fingertip direction and the opposite direction of the x-axis of the image; and take a screenshot from the image based on the fingertip point angle and the position of the fingertip point.
[0180] In some embodiments, the relationship between the screenshot, the angle of the fingertip point, and the position of the fingertip point satisfies any of the following:
[0181] When the angle of the fingertip point is greater than or equal to 0° and less than 75°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is two-thirds of the screenshot size;
[0182] When the angle of the fingertip point is greater than or equal to 75° and less than 100°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the bottom edge of the screenshot is one-quarter of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-half of the screenshot size;
[0183] When the angle of the fingertip point is greater than or equal to 100° and less than 180°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-third of the screenshot size.
[0184] In some embodiments, the apparatus further comprises:
[0185] The trapezoidal correction module is used to perform trapezoidal correction on the screenshot.
[0186] In some embodiments, the apparatus further comprises:
[0187] The processing module is used to perform the following steps when the number of detected regional text boxes is 0:
[0188] Step a: After enlarging the screenshot size, re-execute the screenshot acquisition step;
[0189] Step b: determine whether the number of detected regional text boxes is 0. If so, proceed to step c; if not, do not execute the screenshot acquisition step;
[0190] Step c: Determine whether the number of times the screenshot size has been enlarged reaches a preset number. If not, go to step a. If yes, do not execute the screenshot acquisition step.
[0191] In some embodiments, the area text box is a row text box, and the matching module is specifically used to determine the target row text box that matches the user's fingertip point based on the vertical distance between the user's fingertip point and the four edges of the row text box when the user's fingertip point falls within the row text box; when the user's fingertip point does not fall within the row text box, determine the target row text box that matches the user's fingertip point based on the row text box that intersects with the user's fingertip direction and the vertical distance between the user's fingertip point and the lower edge of the row text box.
[0192] In some embodiments, the matching module is specifically configured to:
[0193] When the line text box is a horizontal text box, if the vertical distance between the user's fingertip point and the four edges of the line text box satisfies: db<(dt+db) / 3 and min(dr,dl)>box_h / 4, then it is determined that the line text box matches the user's fingertip point, wherein db is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl is the vertical distance between the user's fingertip point and the left edge of the line text box, dr is the vertical distance between the user's fingertip point and the right edge of the line text box, and box_h is the height of the line text box;
[0194] When the line text box is a curved text box, if the vertical distance between the user's fingertip and the four edges of the line text box satisfies: db1<(dt1+db1)*2 / 3 and
[0195] min(dr1,dl1)>box_h1 / 4, and abs(c_db)<(dt1+db1) / 3, then it is determined that the line text box matches the user fingertip point, where c_db is the minimum distance from the user fingertip point to the text polygon outline of the curved text box, db1 is the vertical distance between the user fingertip point and the lower edge of the line text box, dt1 is the vertical distance between the user fingertip point and the upper edge of the line text box, dl1 is the vertical distance between the user fingertip point and the left edge of the line text box, and dr1 is the vertical distance between the user fingertip point and the right edge of the line text box. The curved text box satisfies: the ratio of the area of the line text box to the area of the text polygon outline is greater than a preset ratio.
[0196] In some embodiments, the matching module is also used to determine whether the user's fingertip point falls within the line text box when the height of the line text box is greater than a preset first threshold and the vertical distance between the user's fingertip point and the lower edge of the line text box is less than the product of the height of the line text box and a preset first ratio.
[0197] In some embodiments, the matching module is specifically used to determine that the line text box is the target line text box that matches the user's fingertip point if there is a line text box that intersects with the direction of the user's fingertip and the vertical distance between the user's fingertip point and the lower edge of the line text box is less than a preset second threshold, and the second threshold is determined by the average height of the line text boxes in the screenshot.
[0198] In some embodiments, if there is no row text box intersecting with the direction of the user's fingertip, the apparatus further comprises:
[0199] An expansion module, used to expand the size of all text boxes in the screenshot;
[0200] The first judgment module is used to determine whether there is a row text box that intersects with the direction of the user's fingertip. If not, the row text box above the user's fingertip point is determined to be the target row text box that matches the user's fingertip point, and the row text box whose vertical distance between the lower edge and the user's fingertip point is less than the second threshold.
[0201] In some embodiments, the area text box is a line text box, and the apparatus further comprises:
[0202] The second judgment module is used to judge whether the preset expansion conditions are met;
[0203] An image expansion module, configured to expand the size of the screenshot and re-execute the screenshot acquisition step when the image expansion condition is met;
[0204] The expansion conditions include any of the following:
[0205] When no target text box matching the user's fingertip point is obtained, the ratio of the maximum height of all text boxes in the screenshot to the screenshot size is greater than a preset second ratio;
[0206] When a target row text box matching the user's fingertip point is obtained, the ratio of the minimum vertical distance between the target row text box and the edge of the screenshot to the screenshot size is less than a preset third ratio, and the ratio of the height of the target row text box to the screenshot size is greater than a preset second ratio.
[0207] In some embodiments, the regional text box is a line text box, and the detection and recognition module is specifically used to obtain a line text image to be recognized corresponding to the target line text box from the screenshot according to the position information of the target line text box; obtain a character positioning point from the line text image to be recognized according to the direction of the user's fingertip; perform single-word detection on the line text image to be recognized to obtain multiple single-word detection frames, and remove single-word detection frames that meet preset requirements from the multiple single-word detection frames, and the preset requirements are: the width of the single-word detection frame is less than the product of the average width of the single-word detection frame and the preset magnification, and the minimum distance between the single-word detection frame and the edge of the line text image to be recognized is less than a preset third threshold ; The line text image to be recognized is cropped according to the removed single-word detection frame, and text recognition is performed on the cropped line text image to be recognized to obtain a recognized text including multiple characters; if the number of characters included in the recognized text is equal to the number of the single-word detection frames, a target single-word detection frame corresponding to the character positioning point is determined, and a target character among the multiple characters is determined according to the sequence number of the target single-word detection frame in the multiple single-word detection frames; if the number of characters included in the recognized text is not equal to the number of the single-word detection frames, the CTC information in the convolutional recurrent neural network CRNN recognition network is used to estimate the character position, and the target character corresponding to the character positioning point is output.
[0208] In some embodiments, the detection and identification module is specifically configured to, when the target line text box is a horizontal text box, determine an intersection of a first straight line and the fingertip direction when an acute angle formed between the fingertip direction and the lower edge of the target line text box is greater than 45°, wherein the first straight line intersects the target line text box, is parallel to the lower edge, and a distance between the first straight line and the lower edge is one-third of the height of the target line text box; and determine an intersection of the lower edge and the fingertip direction when an acute angle formed between the fingertip direction and the lower edge of the target line text box is less than or equal to 45°;
[0209] When the target line text box is a curved text box, when the fingertip point is located within the character polygon outline of the target line text box, the fingertip point is determined as the intersection point; when the fingertip point is located within the target line text box and outside the character polygon outline, the intersection point of a second straight line and the fingertip direction is determined, where the second straight line is parallel to the lower edge of the target line text box and the distance between the second straight line and the lower edge and the upper edge of the target line text box is half of the height of the target line text box; when the fingertip point is located outside the target line text box, the intersection point of the second straight line and the fingertip direction is determined;
[0210] The coordinates of the intersection are mapped to the line text image to be recognized to obtain the character positioning point.
[0211] In some embodiments, the detection and identification module is further used to add a virtual single-word detection frame between two adjacent single-word detection frames when the distance between the two adjacent single-word detection frames is greater than a preset fourth threshold, and the fourth threshold is determined by the average width of the single-word detection frames.
[0212] In some embodiments, the detection and recognition module is further configured to, when the target line text box is a curved text box, retain S single-word detection boxes closest to the character positioning point if a preset condition is met, and shrink the line text image to be recognized based on the retained S single-word detection boxes, where S is a positive integer;
[0213] The preset conditions include:
[0214] The ratio of the aspect ratio of the to-be-recognized line text image to the number of the single-word detection frames is between 0.9 and 1.8, and the number of the single-word detection frames is greater than S.
[0215] In some embodiments, the detection and recognition module is further configured to convert the character positioning point into an index cross_idx, where the index cross_idx is determined by an x-axis coordinate C_X of the character positioning point in the screenshot and a subsampling number stride of the input image width by the recognition network;
[0216] Decode the index. If the decoded index is the placeholder blank, calculate the difference dist = C_X-cross_idx*stride. If dist is less than 0, take the placeholder blank as the starting point and search for the nearest non-blank character in the recognized text in the order of first searching one character to the left and then searching one character to the right to obtain the target character. If dist is greater than or equal to 0, search for the nearest non-blank character in the recognized text in the order of first searching one character to the right and then searching one character to the left to obtain the target character.
[0217] In some embodiments, the detection and recognition module is further configured to replace the target character with a non-punctuation or space character that is closest to the target character if the target character is a punctuation or space.
[0218] In some embodiments, the detection and recognition module is also used to, if the target character is English, perform forward and backward search on the recognized text starting from the target character until punctuation or space is encountered, to obtain an alternative English character string, segment the alternative English character string, and replace the target character with the segmentation result.
[0219] An embodiment of the present disclosure further provides a character recognition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the program, the character recognition method described above is implemented.
[0220] The embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned character recognition method when executed by a processor.
[0221] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage terminal devices to be detected, or any other non-transmission media that can be used to store information that can be accessed by the computer terminal devices to be detected. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0222] In the various method embodiments of the present disclosure, the serial numbers of the steps cannot be used to limit the order of the steps. For ordinary technicians in this field, without paying any creative work, changes to the order of the steps are also within the scope of protection of the present disclosure.
[0223] It should be noted that the various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, since the embodiments are generally similar to the product embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the product embodiments.
[0224] The above is a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the present disclosure.
Claims
1. A character recognition method, characterized in that: include: A screenshot acquisition step, acquiring a screenshot of the area where the user's fingertip is located; A regional text box detection step, performing text detection on the acquired screenshot to obtain a regional text box, wherein the regional text box includes a row text box and / or a column text box; A matching step, obtaining a target area text box that matches the user's fingertip point; The detection and recognition step performs single-word detection on the target area text box, and determines the target character pointed to by the user's fingertip according to the single-word detection result or according to the single-word detection result and character positioning information.
2. The character recognition method according to claim 1, characterized in that: The screenshot acquisition step includes: Get an image of the user's finger; determining the fingertip points and the finger joint points of the user's finger from the image; Calculating a fingertip direction and a fingertip point angle according to a fingertip point and a finger joint point of a user's finger, wherein the fingertip point angle is an angle between the fingertip direction of the user and a direction opposite to the x-axis of the image; A screenshot is taken from the image according to the fingertip point angle and the fingertip point location.
3. The character recognition method according to claim 2, characterized in that: The relationship between the screenshot, the fingertip angle, and the fingertip position satisfies any of the following: When the fingertip point angle is greater than or equal to 0° and less than 75°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is two-thirds of the screenshot size; When the angle of the fingertip point is greater than or equal to 75° and less than 100°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the size of the screenshot, and the vertical distance between the fingertip point and the left edge of the screenshot is one-half of the size of the screenshot; When the fingertip point angle is greater than or equal to 100° and less than 180°, the fingertip point is located within the screenshot, and the vertical distance between the fingertip point and the lower edge of the screenshot is one-fourth of the screenshot size, and the vertical distance between the fingertip point and the left edge of the screenshot is one-third of the screenshot size.
4. The character recognition method according to claim 1, characterized in that: After the screenshot acquisition step, the method further includes: The step of performing a trapezoidal correction on the screenshot.
5. The character recognition method according to claim 1, characterized in that: The method further comprises: When the number of detected area text boxes is 0, perform the following steps: Step a, after enlarging the size of the screenshot, re-execute the screenshot acquisition step; Step b, determine whether the number of detected regional text boxes is 0, if yes, go to step c; if no, do not execute the screenshot acquisition step; Step c: Determine whether the number of times the screenshot size has been enlarged reaches a preset number. If not, go to step a. If yes, do not execute the screenshot acquisition step.
6. The character recognition method according to claim 1, characterized in that: The area text box is a line text box, and the matching step includes: When the user's fingertip point falls into the line text box, determining a target line text box matching the user's fingertip point according to the vertical distance between the user's fingertip point and four edges of the line text box; When the user's fingertip point does not fall within the line text box, a target line text box matching the user's fingertip point is determined based on the line text box intersecting with the user's fingertip direction and the vertical distance between the user's fingertip point and the lower edge of the line text box.
7. The character recognition method according to claim 6, characterized in that: The step of determining a target line text box matching the user's fingertip point according to the vertical distance between the user's fingertip point and four edges of the line text box comprises: When the line text box is a horizontal text box, if the vertical distance between the user's fingertip point and the four edges of the line text box satisfies: db<(dt+db) / 3 and min(dr,dl)>box_h / 4, then it is determined that the line text box matches the user's fingertip point, wherein db is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl is the vertical distance between the user's fingertip point and the left edge of the line text box, dr is the vertical distance between the user's fingertip point and the right edge of the line text box, and box_h is the height of the line text box; When the line text box is a curved text box, if the user's fingertip point is in contact with the line text box, The vertical distance between the edges satisfies: db1<(dt1+db1)*2 / 3 and min(dr1,dl1)>box_h1 / 4, and abs(c_db)<(dt1+db1) / 3, then it is determined that the line text box matches the user's fingertip point, wherein c_db is the minimum distance from the user's fingertip point to the text polygon outline of the curved text box, db1 is the vertical distance between the user's fingertip point and the lower edge of the line text box, dt1 is the vertical distance between the user's fingertip point and the upper edge of the line text box, dl1 is the vertical distance between the user's fingertip point and the left edge of the line text box, dr1 is the vertical distance between the user's fingertip point and the right edge of the line text box, and the curved text box satisfies: the ratio of the area of the line text box to the area of the text polygon outline is greater than a preset ratio.
8. The character recognition method according to claim 6, characterized in that: The method further comprises: When the height of the line text box is greater than a preset first threshold and the vertical distance between the user's fingertip point and the lower edge of the line text box is less than the product of the height of the line text box and a preset first ratio, it is determined that the user's fingertip point falls into the line text box.
9. The character recognition method according to claim 6, characterized in that: The step of determining a target line text box matching the user's fingertip point according to the line text box intersecting with the user's fingertip direction and the vertical distance between the user's fingertip point and the lower edge of the line text box comprises: If there is a line text box that intersects with the direction of the user's fingertip, and the vertical distance between the user's fingertip point and the lower edge of the line text box is less than a preset second threshold, the line text box is determined to be the target line text box that matches the user's fingertip point, and the second threshold is determined by the average height of the line text boxes in the screenshot.
10. The character recognition method according to claim 9, characterized in that: If there is no row text box intersecting with the direction of the user's fingertip, the method further includes: Enlarging the size of all text boxes in the screenshot; Determine whether there is a row text box that intersects with the direction of the user's fingertip. If not, determine among the row text boxes above the user's fingertip point that the row text box whose vertical distance between the lower edge and the user's fingertip point is less than the second threshold is the target row text box that matches the user's fingertip point.
11. The character recognition method according to claim 1, characterized in that: The area text box is a line text box. After the matching step, the method further includes: A judgment step, judging whether a preset image expansion condition is met; an image enlargement step, when the image enlargement condition is met, the screenshot size is enlarged and the screenshot acquisition step is re-executed; The expansion conditions include any of the following: When the target line text box matching the user's fingertip point is not obtained, the ratio of the maximum height of all line text boxes in the screenshot to the screenshot size is greater than a preset second ratio; When a target row text box matching the user's fingertip point is obtained, the ratio of the minimum vertical distance between the target row text box and the edge of the screenshot to the screenshot size is less than a preset third ratio, and the ratio of the height of the target row text box to the screenshot size is greater than a preset second ratio.
12. The character recognition method according to claim 1, characterized in that: The area text box is a line text box, and the detection and identification step includes: Acquire a text line image to be identified corresponding to the target text line box from the screenshot according to the position information of the target text line box; Obtaining a character positioning point from the to-be-recognized line text image according to the fingertip direction of the user; Performing single-word detection on the line text image to be recognized to obtain a plurality of single-word detection frames, and removing the single-word detection frames that meet preset requirements from the plurality of single-word detection frames, wherein the preset requirements are: the width of the single-word detection frame is less than the product of the average width of the single-word detection frame and a preset magnification, and the minimum distance between the single-word detection frame and the edge of the line text image to be recognized is less than a preset third threshold; Cropping the to-be-recognized line text image according to the removed single-word detection frame, and performing text recognition on the cropped to-be-recognized line text image to obtain a recognized text including a plurality of characters; If the number of characters included in the recognized text is equal to the number of the single-word detection frames, determining a target single-word detection frame corresponding to the character positioning point, and determining a target character among the multiple characters according to the sequence number of the target single-word detection frame in the multiple single-word detection frames; If the number of characters included in the recognized text is not equal to the number of the single-word detection boxes, the connected text classification CTC information in the convolutional recurrent neural network CRNN recognition network is used to estimate the character position, and the target character corresponding to the character positioning point is output.
13. The character recognition method according to claim 12, characterized in that: The step of obtaining a character positioning point from the to-be-recognized line text image according to the fingertip direction of the user comprises: When the target line text box is a horizontal text box, when the acute angle formed by the fingertip direction and the lower edge of the target line text box is greater than 45°, determine the intersection of a first straight line and the fingertip direction, the first straight line intersecting the target line text box, being parallel to the lower edge, and the distance between the first straight line and the lower edge being one third of the height of the target line text box; when the acute angle formed by the fingertip direction and the lower edge of the target line text box is less than or equal to 45°, determine the intersection of the lower edge and the fingertip direction; When the target line text box is a curved text box, when the fingertip point is located within the character polygon outline of the target line text box, the fingertip point is determined to be an intersection point; when the fingertip point is located within the target line text box and outside the character polygon outline, the intersection point of a second straight line and the fingertip direction is determined, the second straight line is parallel to the lower edge of the target line text box and the distance between the second straight line and the lower edge and the upper edge of the target line text box is half of the height of the target line text box; when the fingertip point is located outside the target line text box, the intersection point of the second straight line and the fingertip direction is determined; The coordinates of the intersection are mapped to the line text image to be recognized to obtain the character positioning point.
14. The character recognition method according to claim 12, characterized in that: After obtaining multiple single-word detection frames, the method further includes: When the distance between two adjacent single-word detection frames is greater than a preset fourth threshold, a virtual single-word detection frame is added between the two adjacent single-word detection frames, and the fourth threshold is determined by the average width of the single-word detection frames.
15. The character recognition method according to claim 12, characterized in that: After obtaining multiple single-word detection frames, the method further includes: When the target line text box is a curved text box, if a preset condition is met, S single-word detection boxes closest to the character positioning point are retained, and the line text image to be recognized is shrunk according to the retained S single-word detection boxes, where S is a positive integer; The preset conditions include: The ratio of the aspect ratio of the to-be-recognized line text image to the number of the single-word detection frames is between 0.9 and 1.8, and the number of the single-word detection frames is greater than S.
16. The character recognition method according to claim 12, characterized in that: The method further comprises: If the target character is a punctuation mark or a space, the target character is replaced with the character that is the nearest non-punctuation mark or space mark to the target character.
17. The character recognition method according to claim 12, characterized in that: The method further comprises: If the target character is in English, the recognized text is searched forward and backward starting from the target character until a punctuation mark or a space is encountered, to obtain an alternative English character string, segment the alternative English character string, and replace the target character with the segmentation result.
18. A character recognition device, characterized in that: include: A screenshot acquisition module is used to obtain a screenshot of the area where the user's fingertip is located; A line text box detection module, used to perform text detection on the acquired screenshot to obtain a line text box; A matching module, used for acquiring a target line text box matching the user's fingertip point; The detection and recognition module is used to perform single-word detection on the target line text box, and determine the target character pointed to by the user's fingertip according to the single-word detection result or according to the single-word detection result and character positioning information.
19. A character recognition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that: When the processor executes the program, the character recognition method according to any one of claims 1 to 17 is implemented.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the character recognition method as described in any one of claims 1 to 17 are implemented.