Text localization method and apparatus, and device and storage medium
By detecting the position of fingertips, identifying and filtering out the paragraph text box with the largest area, the problem of fixed position of the camera and horizontal placement of text is solved, and accurate positioning and convenient identification of text is achieved.
Patent Information
- Application Number
- PCT/CN2024/141136
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, when a camera takes a text image, a fixed position and angle are required, and the text needs to be placed horizontally, otherwise the content of the text to be identified cannot be accurately positioned.
By detecting the position of the fingertips on the text page, determine the area image of the text to be identified, identify the first text box with the largest area in the paragraph text box, and filter out the paragraph with the center point of the paragraph in the target text box as the text to be identified.
It realizes accurate positioning of the text to be identified without limiting the camera position and angle, reducing the requirements for the text placement angle and improving the accuracy and convenience of positioning.
Smart Images

Figure CN2024141136_03072025_PF_FP_ABST
Abstract
Description
Text positioning method, device, equipment and storage medium Technical Field
[0001] The present application relates to the field of information recognition technology, and in particular to a text positioning method, apparatus, device and storage medium. Background Art
[0002] In daily life and work, it is often necessary to identify and extract text. For example, when a part of the content in a text page needs to be translated, or when a part of the content in a text page needs to be copied, it is necessary to identify the part of the text content and extract the identified part of the text content. In the related art, the text to be identified is photographed by a camera, and the photographed text image is then cropped to obtain an image of the area where text recognition is required. However, when photographing the text to be identified, the position and shooting angle of the camera are fixed, and the text to be identified needs to be placed horizontally. Otherwise, the text in the captured text image is tilted, making it impossible to accurately locate the text content to be identified. Summary of the Invention
[0003] The purpose of this application is to provide a text positioning method, device, equipment and storage medium to overcome the defects that when shooting a text image containing text content to be identified, the camera and shooting angle need to be fixed, and the text to be identified needs to be placed completely horizontally, otherwise the text content to be identified cannot be accurately located.
[0004] The technical solutions adopted by this application to solve the technical problems are as follows:
[0005] In a first aspect, this embodiment discloses a text positioning method, wherein the method includes:
[0006] detecting a fingertip position of a fingertip on a text page, and determining a region image of text to be recognized in the text page based on the fingertip position;
[0007] Identify a first text box with the largest area among the paragraph text boxes corresponding to the paragraphs in the regional image;
[0008] determining a target text box according to the first text box and the fingertip position;
[0009] One or more paragraphs whose center points are located within the target text frame are screened out, and the text contained in each of the screened out paragraphs is determined as the text to be recognized.
[0010] Optionally, the fingertip position includes a first fingertip position located at the upper left point of the text page content and a second fingertip position located at the lower right point of the text page content; and the step of detecting the fingertip position of the fingertip on the text and determining the regional image of the text to be recognized based on the fingertip position includes:
[0011] Acquire a photographed image of a fingertip on a text page, and identify a first fingertip position and a second fingertip position in the photographed image;
[0012] Calculating a midpoint of a connecting line between the first fingertip position and the second fingertip position and a length of the connecting line;
[0013] The midpoint of the connecting line is used as the center of the cut-off image, and the length of the connecting line is used as the length and width of the cut-off image. An image is cut off from the captured image, and the cut-off image is determined as the area image of the text to be recognized.
[0014] Optionally, the step of acquiring a photographed image of a fingertip on a text page and identifying a first fingertip position and a second fingertip position in the photographed image includes:
[0015] photographic image of the tip of a finger positioned on a page of text;
[0016] Locating a hand region frame at the position of the hand from the captured image;
[0017] cropping the hand region frame from the captured image to obtain a hand image;
[0018] Key point detection is performed on the hand image to obtain a detected first fingertip position and a second fingertip position.
[0019] Optionally, after the step of performing key point detection on the hand image to obtain the detected first fingertip position and the second fingertip position, the method further includes:
[0020] Detect whether the camera's viewing axis has changed;
[0021] When it is detected that the visual axis of the camera changes, an updated image after the visual axis changes is obtained;
[0022] extracting key points and feature descriptors from the hand image and the update image respectively;
[0023] Matching the hand image and the updated image using the key points and feature descriptors to obtain a transformation matrix between the hand image and the updated image;
[0024] The first fingertip position in the updated image is calculated using the transformation matrix, and the first fingertip position in the updated image is used as the updated first fingertip position.
[0025] Optionally, the step of identifying a first text box with the largest area among the paragraph text boxes corresponding to the paragraphs in the regional image includes:
[0026] Inputting the region image into a trained text detection model to obtain a second text box corresponding to each line of text in the region image output by the text detection model;
[0027] According to the geometric position relationship between the second text boxes corresponding to each text line, the paragraphs to which each text line belongs are identified, and then the paragraph text boxes corresponding to each paragraph are obtained;
[0028] Filter out the first text box with the largest area from the paragraph text boxes corresponding to each paragraph.
[0029] Optionally, the fingertip positions include a first fingertip position located at an upper left point of the text page content and a second fingertip position located at a lower right point of the text page content;
[0030] The step of selecting the first text box with the largest area from the paragraph text boxes corresponding to each paragraph includes:
[0031] Obtaining a segment intersecting with a line connecting the fingertips of the first fingertip position and the second fingertip position;
[0032] Filter out the first text box with the largest area from the paragraph text boxes corresponding to the intersecting paragraphs.
[0033] Optionally, the step of determining a target text box according to the first text box and the fingertip position includes:
[0034] Get the long side direction of the first text box;
[0035] A target text box is determined according to the long side direction of the first text box and the fingertip position.
[0036] Optionally, the step of filtering out one or more paragraphs whose center points are located within the target text box includes:
[0037] Get the four corner coordinates of the paragraph text box corresponding to each paragraph in the target text box;
[0038] Calculate the center point of each paragraph text box according to the four corner coordinates of each paragraph text box;
[0039] According to the text boxes corresponding to the respective paragraphs, it is determined in sequence whether the center point of the paragraph text box corresponding to the respective paragraphs is located within the target text box;
[0040] It is determined in sequence whether the center point of the paragraph text box corresponding to each paragraph is located in the target text box, and one or more paragraphs whose center points are located in the target text box are obtained.
[0041] In a second aspect, this embodiment discloses a text positioning device, which includes:
[0042] a region image determination module, configured to detect a fingertip position of a finger on a text page, and determine a region image of text to be recognized in the text page based on the fingertip position;
[0043] A text box identification module is used to identify a first text box with the largest area among the paragraph text boxes corresponding to each paragraph in the regional image;
[0044] a target text box identification module, configured to determine a target text box according to the first text box and the fingertip position;
[0045] The text area determination module is used to determine each paragraph whose center point is located in the target text box as the paragraph where the text to be recognized is located, and locate the text to be recognized.
[0046] In a third aspect, this embodiment further provides an intelligent device, which includes a memory, a processor, and a text positioning program stored in the memory and executable on the processor. When the processor executes the text positioning program, the steps of the text positioning method are implemented.
[0047] In a fourth aspect, this embodiment further discloses a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the text positioning. Beneficial effects:
[0048] The present embodiment discloses a text positioning method, apparatus, device and storage medium, which detects the fingertip position of a finger on a text page, determines the regional image of the text to be identified based on the fingertip position; identifies the first text box with the largest area among the paragraph text boxes corresponding to each paragraph in the regional image; determines the target text box based on the first text box, filters out one or more paragraphs whose paragraph center points are within the target text box, and determines the text contained in each filtered paragraph as the text to be identified. In the embodiment of the present application, there is no need to limit the fixed position and angle of the camera for shooting the image of the text to be identified, nor is there any need to limit the text to be identified to be placed completely horizontally when shooting. The text to be identified can be accurately positioned based on the position of the fingertip, so it is easy to implement and has a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a flowchart of the steps of the text positioning method provided in this embodiment;
[0050] FIG2 is a schematic diagram of positioning using a finger provided in this embodiment;
[0051] FIG3 is a schematic diagram of a region image cut out from a text image provided by this embodiment;
[0052] FIG4 is a schematic diagram of merging regional images into paragraphs provided by this embodiment;
[0053] FIG5 is a schematic diagram of filtering out the text box with the largest area in a regional image provided by this embodiment;
[0054] FIG6 is a schematic diagram of locating the text area to be recognized provided by this embodiment
[0055] FIG7 is an example diagram of a text area to be identified by locating the area provided by this embodiment;
[0056] FIG8 is a block diagram of the principle structure of the text positioning device provided in this embodiment;
[0057] FIG9 is a block diagram of the principle structure of the smart device in an embodiment of the present application. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0059] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0060] In daily work and life, we often encounter the use of smart devices to identify and extract information in text to obtain characters in paper documents.
[0061] The method currently used to identify or extract characters from paper documents is generally the OCR (Optical Character Recognition) text recognition method. OCR text recognition refers to using a scanner or digital camera to scan or photograph a paper document to obtain an image of the text to be extracted. After binarization, noise removal, and tilt correction of the image, character segmentation is performed, and then character recognition is performed to locate the area where the text to be recognized is located.
[0062] When using a digital camera to photograph a paper document, it's necessary to fix the camera's position and angle, and lay the document flat, in order to identify the area where text recognition is needed. However, if the camera moves, the angle of the camera isn't directly facing the document, or the document is tilted relative to the camera, the characters contained in the document won't be accurately recognized. If the characters to be recognized and extracted are part of a page on a paper document, accurately locating the text becomes even more difficult.
[0063] In order to overcome the above technical problems, this embodiment provides a text positioning method, which obtains the fingertip position of the fingertip on the text page, and then obtains the regional image containing the text to be identified based on the fingertip position, analyzes the text lines in the regional image based on the fingertip position, obtains the paragraph text box corresponding to the paragraph, and then determines the paragraph text box with the largest area, and finally determines all paragraphs whose center points are located in the paragraph text box with the largest area as the text to be identified, thereby realizing the positioning of the text to be identified. The method provided by this embodiment is first based on the accurate positioning of the fingertip position, and then analyzes the various text lines and paragraphs that appear between the fingertip positions in the text file based on the fingertip position, which can realize the accurate analysis and positioning of the text lines between the fingertip positions, improves the positioning accuracy, and thus reduces the requirements for the camera shooting position and shooting angle, as well as the placement angle of the paper text, providing a guarantee for more convenient positioning of the text to be identified.
[0064] The method, apparatus and device provided in this embodiment will be further described in detail below with reference to the accompanying drawings.
[0065] As shown in FIG1 , a text positioning method provided by this embodiment includes:
[0066] Step S1: Detect the position of a fingertip on a text page, and determine a region image of the text to be recognized based on the fingertip position.
[0067] When it is necessary to locate the text to be recognized, the user first places the fingertips on both sides of the text to be recognized, thereby locking the text to be recognized.
[0068] In a specific implementation, if automatic detection is configured using a camera installed on a smart device, when the camera detects that the fingertips of both hands are placed on both sides of the text to be recognized, an image of the fingertips on the text page is captured, and the two fingertip positions in the captured image are identified to obtain a first fingertip position and a second fingertip position, respectively. The camera can also capture two images of the two fingertips at the upper left and lower right points of the text page content in two separate times to further recognize the first fingertip position and the second fingertip position.
[0069] In addition, when the smart device is a portable device, such as when taking pictures with a mobile phone, a possible way to achieve this is to take a picture of the fingertips of two hands placed on the text with the assistance of other people, or to take two images in succession, with the first picture containing the first fingertip position located at the upper left point of the text page content, and the second picture containing the second fingertip position located at the lower right point of the text page content.
[0070] As shown in Figure 2, if the fingertip of the left hand is placed at the upper left point of the text to be recognized, and the fingertip of the right hand is placed at the lower right point of the text to be recognized, the paragraph text between the upper left point and the lower right point belongs to the text to be recognized.
[0071] In specific implementation, as shown in Figure 2, two hands can be used to lock the text to be recognized. If one hand is used to lock the text to be recognized, it is necessary to identify the fingertip position of the upper left point and the fingertip position of the lower right point twice respectively. Specifically, first identify the position of the fingertip at the upper left point, and then identify the position of the fingertip at the lower right point within a preset time, thereby achieving the locking of the text to be recognized using one hand.
[0072] Furthermore, the fingertip position can be identified by obtaining an image through a camera. In order to obtain a more accurate fingertip position, a video of the fingertip placed on the text page can also be captured by a camera, and each frame of the video can be analyzed to achieve this.
[0073] Specifically, if the fingertip position is identified by an image captured by a camera, it is necessary to first obtain an image captured by the camera containing the finger as shown in Figure 2, and then analyze the captured image to obtain the fingertip position. The steps include:
[0074] Step S11: Acquire a photographed image of a fingertip on a text page, and identify a first fingertip position and a second fingertip position in the photographed image.
[0075] The method for obtaining the captured image can be to capture the image using a camera of the smart device, or, if the smart device is not equipped with a camera, to obtain a captured image containing the fingertip locked on the text to be recognized from another smart device. The smart device can be a portable device equipped with a camera or a near-eye display device, such as a mobile phone, tablet, laptop, AR glasses, or XR helmet.
[0076] After the captured image is acquired, the first fingertip position and the second fingertip position of the finger in the captured image are identified from the captured image, as shown in FIG2 , which are positions P1 and P2 in the figure.
[0077] Furthermore, the step of acquiring the captured image and identifying the first fingertip position and the second fingertip position in the captured image in this step includes:
[0078] An image is captured in which a fingertip is located on a text page, and a hand region frame containing finger position information is extracted from the captured image.
[0079] Use a camera to capture an image, or receive an image from another smart device. Extract information from the captured image to obtain a hand region frame containing the finger positions. The hand region frame is a graphic that marks the identified hand position with a square frame.
[0080] Specifically, the hand region frame is extracted from the captured image to obtain finger position information, and the hand image is cropped from the captured image based on the hand region frame.
[0081] Key point detection is performed on the hand image to obtain detected fingertip positions.
[0082] The cropped hand image is subjected to key point detection to obtain the detected first fingertip position and second fingertip position. Specifically, the method for performing key point detection on the hand image can be implemented using a key point detection algorithm based on target detection, or it can be implemented using a trained key point detection model. Commonly used key point detection algorithms include: corner detection. The key point detection model is based on the use of hand key point features to input an image containing hand feature annotations into a trained hand key point detection model. The hand key point detection model includes a convolutional network and a fully connected layer. The input hand image is processed by the convolutional network and the fully connected layer to output the coordinates of the hand key points.
[0083] In one embodiment, a trained hand detection model may be used to detect the hand position in the captured image.
[0084] Specifically, the hand detection model is obtained by training a preset network model using a sample set. The sample set includes a positive sample set and a negative sample set. The sample images contained in the positive sample set are images of the fingertips of the collected fingers placed on a paper text page, and the image is marked with a hand location area frame of the fingertips. The sample images contained in the negative sample set do not contain images of hands, so the hand location area frame is not marked. The sample images in the sample set are input into the preset network model, and the preset network model is trained to obtain a trained hand detection model.
[0085] When a captured image corresponding to the text to be recognized is obtained, the captured image is input into the trained hand detection model to obtain a hand area frame containing hand position information output by the hand detection model.
[0086] Furthermore, when a video of a fingertip placed on a paper text page is obtained by a camera, the coordinates of the fingertip position in each frame of the video are recorded. Then, in the consecutive preset frame images, it is determined whether the range of change of the horizontal and vertical coordinates of the fingertip position exceeds the preset pixel threshold. If it does not exceed, it is determined that the fingertip position in the video remains stationary, and the fingertip position that remains stationary in the current video frame is used as the detected fingertip position. For example: If the video stream contains 30 frames of images per second, if the range of change of the horizontal and vertical coordinates of the fingertip position does not exceed 20 pixels in 60 consecutive frames of images, it is considered to be stable for 2 seconds, and the fingertip coordinates of the currently stable upper left point and lower right point are determined as the coordinates of the detected fingertip position.
[0087] If the operation of locking the hand on the text to be recognized is completed by one hand, it is necessary to first determine the coordinates of one of the fingertip positions, and then determine the coordinates of the other fingertip position within the next preset time (which may be within a few seconds), so as to achieve the determination of the corresponding coordinates of the two fingertip positions. Specifically, when the fingertip of the hand is stable at the upper left point of the text page content, an image of the fingertip when it is stable is captured, and the captured image is input into the hand detection model to obtain the hand area frame output by the hand detection model when the hand is at the upper left point. In the following time, when it is detected that the fingertip of the hand is stable at the lower right point of the text page content, the current image is captured, and the image of the current fingertip at the lower right point is input into the hand detection model to obtain the hand area frame output by the hand detection model with the finger position at the lower right point, thereby obtaining the hand area frame position coordinates located at the upper left point and the lower right point respectively. Then, based on the hand area frame corresponding to the upper left point and the hand area frame corresponding to the lower right point, the left and right hand area frames are cropped to obtain the hand images corresponding to the left and right sides. The hand images are then subjected to key point detection respectively to obtain the detected fingertip positions of the upper left point and the lower right point.
[0088] In this step, a captured image is first obtained, the hand position is located from the captured image, and a hand image is obtained by cropping the hand image. The fingertip positions in the hand image are then detected using a hand detection model to accurately determine the fingertip positions. In this embodiment, the fingertip positions are detected more accurately because the fingertip positions are located step by step.
[0089] When the smart device is a smart wearable device, such as AR glasses, changes in the user's position or wearing angle will cause the camera's visual axis to change. Alternatively, although the camera of the smart device is fixed, the position of the camera on the smart device may change due to external factors, which will also cause the camera's visual axis to change. As the camera's visual axis changes, the position of the first fingertip at the upper left point changes. In order to obtain a more accurate first fingertip position, this step needs to correct the first fingertip position. Therefore, this step also includes:
[0090] Step S111 , detecting whether the visual axis of the camera has changed; when it is detected that the visual axis of the camera has changed, acquiring an updated image after the visual axis has changed.
[0091] When the visual axis of the camera changes, there will be differences in the images captured successively after the visual axis changes. Therefore, whether the visual axis of the camera has changed can be determined based on whether there are differences in the captured images.
[0092] Specifically, the camera captures images at preset intervals and calculates the pixel difference between two consecutive images. Whether there is a difference between the two images is determined based on whether the pixel difference is greater than a preset pixel difference value. If a difference is determined between the two images, it is determined that the camera's visual axis has changed. Therefore, in actual use, after capturing the current image, the camera captures another image at a preset time. The pixel difference between the two images is calculated. If the pixel difference exceeds the preset pixel difference value, it is determined that the camera's visual axis has changed.
[0093] For example: after capturing the first image of the current fingertip on the text page, if the preset time is 3 seconds, the second image is captured 3 seconds later. When the pixel difference between the first image and the second image exceeds the preset pixel difference value, it indicates that there is a difference between the first image and the second image, and it is determined that the camera's visual axis has changed.
[0094] In specific implementation, when the camera is fixed and the camera's visual axis is less likely to change due to external reasons, it can be set not to detect the change in the camera's visual axis and not to correct the first fingertip position. The correction of the first fingertip position is triggered only when it is detected that the position of the smart device itself has changed or the camera's shooting parameters have changed. For example: when the smart device is a computer fixed on the wall, the position of the camera on the computer is fixed, and its visual axis generally does not change, so there is no need to correct the first fingertip position. Only when the camera is not fixed, it is necessary to detect whether the visual axis has changed. When the smart device is a mobile phone or a tablet, although the cameras of these portable smart devices can be moved, their positions are not moving in real time when in use. Therefore, the first fingertip position can be corrected when a change in the visual axis is detected. In addition, when the smart device is a wearable smart device, it can also be set not to detect whether the visual axis has changed, and it can be set to automatically obtain an updated image at preset time intervals when the user is in use, and update the first fingertip position. Because the position of the camera on the wearable smart device changes in real time during use, during use, it can be directly set to update the first fingertip position at preset time intervals without having to detect whether the visual axis has changed.
[0095] Since the change of the visual axis of the camera may cause the position of the first fingertip to deviate, it is necessary to capture an updated image after the visual axis changes, and correct the first fingertip position according to the updated image.
[0096] Step S112: extract key points and feature descriptors from the hand image and the updated image respectively; use the key points and feature descriptors to match the hand image and the updated image to obtain a transformation matrix between the hand image and the updated image.
[0097] Extract key points and feature descriptors of the hand image and the update image, perform feature matching, and find the transformation matrix between the hand image and the update image.
[0098] Furthermore, the coordinates of the first fingertip position in the hand image can be obtained, and an n*n area centered on the coordinates can be cropped as the first image, where the value of n can be between 1 / 4 and 1 of the length or width of the text page. When the screen moves, an updated image is captured, and an n*n area is cropped in the current frame as the second image based on the first fingertip position in the updated image. By extracting key points and feature descriptors from the first image and the second image respectively, the key points and feature descriptors are used to match the first image and the second image to obtain the transformation matrix between the first image and the second image, thereby simplifying the process of calculating the transformation matrix.
[0099] Specifically, using key points and feature descriptors to match the first image and the second image is essentially to use key points and feature descriptors to calculate the distance between each key point. A key point is selected in the first image, and then the distance values between the selected key point and each key point in the second image are calculated in sequence (that is, the distance values between the feature descriptor corresponding to the selected key point and the feature descriptor corresponding to each key point in the second image are calculated. The feature descriptor includes not only the key point but also the pixels around the key point that contribute to it. The matched key points have similar feature descriptors). The point with the smallest distance value is returned, and the key point with the smallest distance value in the second image is used as the matched key point, thereby obtaining the matching relationship between each key point in the first image and each key point in the second image. Common distance calculation methods include: Euclidean distance, Hamming distance, cosine distance, etc. The transformation matrix is obtained based on the matching relationship between the key points between the first image and the second image.
[0100] Specifically, a key point refers to the location of a feature point in an image, which has information such as direction and scale. A feature point is a point with a relatively special location in the image, such as a corner point or a point on an edge. The descriptor is usually a vector that describes the pixel information in the neighborhood of the key point. Multiple key points are selected in the hand image and the updated image respectively, and feature descriptors are calculated based on the selected key points. The feature descriptor can be a binary descriptor. For example, when the distance between two other points and a key point is far, if it is far, it is 1, otherwise it is 0. After obtaining the key points and feature descriptors in the hand image and the updated image respectively, the key points and feature descriptors are matched, and the positional relationship between the hand image and the updated image is calculated to obtain the transformation matrix between the hand image and the updated image.
[0101] Step S113: Calculate the first fingertip position in the updated image using the transformation matrix, and use the first fingertip position in the updated image as the updated first fingertip position.
[0102] According to the calculated transformation matrix, the first fingertip position at the upper left point in the hand image is perspective transformed through the obtained transformation matrix to obtain its position coordinates on the updated image, thus obtaining the position coordinates corresponding to the corrected first fingertip position at the upper left point.
[0103] Step S12: Calculate the position of the midpoint of the connecting line between the first fingertip position and the second fingertip position, and the length of the connecting line between the first fingertip position and the second fingertip position.
[0104] After the first fingertip position and the second fingertip position are detected in step S11, the midpoint of the connecting line between the two coordinate points and the length of the connecting line between the two coordinate points are calculated based on the located position coordinates. Specifically, the midpoint coordinates of the connecting line and the length of the connecting line can be directly calculated based on the located coordinate values.
[0105] Step S13 : Taking the midpoint of the connecting line as the center point and the length of the connecting line as the length and width, a cutout image is captured from the captured image, and the cutout image is determined as the area image of the text to be recognized.
[0106] The midpoint of the connecting line is used as the center point, and the length of the connecting line is used as the length and width of the intercepted area image. The area image of the area where the text to be recognized is located is determined from the captured image.
[0107] Step S2: Identify the first text box with the largest area among the paragraph text boxes corresponding to the paragraphs in the regional image.
[0108] When the regional image of the area where the text to be identified is located is determined in the above step S1, the regional image is analyzed to obtain the text box corresponding to each text line in the regional image, and then the positional relationship between the text boxes corresponding to two adjacent text lines is calculated to determine whether the two adjacent text lines belong to the same paragraph. Specifically, it can be determined by judging whether there is a blank line between the two adjacent text lines. If there is a blank line between the two adjacent text lines, they belong to different paragraphs. If there is no blank line between the two adjacent text lines, they belong to the same paragraph. If they belong to the same paragraph, they are divided into the same paragraph. Otherwise, the text line is divided into another paragraph. The distance between the text boxes corresponding to multiple adjacent text lines is used for judgment. Each text line is divided and merged into multiple paragraphs to obtain multiple paragraphs. The area of the paragraph is then judged, and the paragraph text box corresponding to the paragraph with the largest area is selected.
[0109] Specifically, the step of identifying the first text box with the largest area includes:
[0110] Step S21: input the region image into a trained text detection model to obtain a second text box corresponding to each line of text in the region image output by the text detection model.
[0111] The trained text detection model is used to identify each text line in the region image. Specifically, a text line is a text box corresponding to a single line of text, which contains a line of text or characters. Users usually edit the content of the entire document page by editing each text line.
[0112] Specifically, the text detection model in this step can be implemented using an open source text detection model or by calling an OCR recognition API interface.
[0113] Step S22: Identify the paragraphs where each text line is located based on the geometric positional relationship between the second text boxes corresponding to each text line, and then obtain the paragraph text boxes corresponding to each paragraph.
[0114] When the second text boxes corresponding to each text line in the regional image are identified, the geometric position relationship between each second text box is obtained respectively (the geometric position relationship includes: distance relationship, orientation relationship, overlap, and whether the left and right are aligned, etc.), and the text lines are merged according to the identified geometric position relationship to obtain the merged paragraphs, and then the paragraph text boxes corresponding to each paragraph are obtained.
[0115] Specifically, the text box of a text line is composed of four points located at the four corners. First, the position coordinates of the points at the four corners of each text line are obtained. Based on the position coordinates of the four corners, the distance, orientation relationship, overlap, left-right alignment, and other indicators between the corresponding second text boxes of adjacent text lines are calculated. Then, based on the above calculation results, the positional relationship between the text boxes of two adjacent text lines is determined. If the positional relationship satisfies the requirement of belonging to the same paragraph, the two text lines are merged into the same paragraph. Otherwise, the two adjacent text lines are divided into two different paragraphs.
[0116] Step S23: Filter out the first text box with the largest area from the paragraph text boxes corresponding to each paragraph.
[0117] The position coordinates of the four corners of the paragraph text box corresponding to each paragraph are obtained, and the area of the paragraph text box corresponding to each paragraph is calculated based on the position coordinates of the four corners, and then the first text box with the largest area is selected.
[0118] In order to more accurately filter out the first text box with the largest area from the paragraph text boxes corresponding to each paragraph in the text to be recognized, the coordinates of the located fingertip position can also be used to establish a connecting line between the two fingertip positions, and the first text box with the largest area can be filtered out from the paragraph intersecting with the connecting line.
[0119] Specifically, the fingertip position includes a first fingertip position located at the upper left point of the text page content and a second fingertip position located at the lower right point of the text page content. This step also includes: obtaining paragraphs that intersect with the connecting line of the fingertips of the first fingertip position and the second fingertip position, and filtering out the first text box with the largest area from the intersecting paragraphs.
[0120] Step S3: determining a target text box according to the first text box and the fingertip position.
[0121] Specifically, in order to more accurately locate the area where the text to be recognized is located, this step includes:
[0122] The long side direction of the first text box is obtained; and a target text box is determined according to the long side direction of the first text box and the fingertip position.
[0123] After determining the first text box with the largest area, the two fingertip points P1 and P2, as well as the direction of the long side of the first text box with the largest area, yield a large target text box (P1EP2F, shown in Figure 6) and the blue box in Figure 7. Since the positions of the two fingertips (P1 and P2) and the direction of the long side of the first text box have been calculated, the rectangular area corresponding to the target text box can be determined based on the direction of the long side and the positions of the two fingertips.
[0124] In this embodiment, the target text box is determined to be the text box corresponding to the area locked between the two fingertips on the text page, thereby locating the text to be recognized. Since this embodiment selects the long side direction of the first text box with the largest area among the paragraph text boxes as the direction of one side of the target text box, this not only reduces the accuracy requirement for directly locating the area of the document to be recognized, but also makes it easier to implement.
[0125] Step S4: Filter out one or more paragraphs whose center points are located within the target text box, and determine the text contained in each of the filtered paragraphs as the text to be recognized.
[0126] After obtaining the paragraph text boxes corresponding to multiple paragraphs in the above steps, the paragraph center points of the paragraph text boxes corresponding to each paragraph are obtained in turn, and it is determined one by one whether the paragraph center points of the paragraph text boxes corresponding to each paragraph are located in the target text box determined above. If the paragraph center point is located in the target text box, it is determined that the text content corresponding to the paragraph belongs to the text to be recognized. If the paragraph center point is not located in the target text box, it is determined that the text content corresponding to the paragraph does not belong to the text to be recognized. All paragraphs whose corresponding text contents are determined to belong to the text to be recognized are merged together to form the area where the text to be recognized is located to be located in this embodiment.
[0127] Step S41: The coordinates of the four corners of each paragraph text box corresponding to each paragraph in the target text box are calculated, and then the center point of each paragraph text box is calculated according to the coordinates of the four corners.
[0128] When the paragraph text boxes corresponding to each paragraph are obtained in step S2, the coordinates of the four corners of the paragraph text boxes corresponding to each paragraph in the target text box are obtained respectively. According to the obtained coordinates of the four corners of the paragraph text boxes corresponding to each paragraph, the paragraph center points of the paragraph text boxes corresponding to each paragraph are calculated in sequence.
[0129] Step S42: determine in sequence whether the center point of the paragraph text box corresponding to each paragraph is located in the target text box, and obtain one or more paragraphs whose center points are located in the target text box.
[0130] According to the coordinates of the four corners of the paragraph text box corresponding to each paragraph, it is determined in turn whether the center point of the paragraph text box of each paragraph is located in the target text box.
[0131] Specifically, as shown in Figures 6 and 7, if the center point K of the text box corresponding to a paragraph is located inside the target text box P1EP2F, then the paragraph is the paragraph corresponding to the text to be recognized. Specifically, the four corner points of the target text box are arranged in the order of P1EP2F in a clockwise direction. The vector products of the lines connecting the center point of each paragraph with the four corner points of P1EP2F and the sides of the text box connected to them are calculated respectively, and each vector product is determined to be greater than 0. If each vector product is greater than 0, it means that the center point is located inside the target text box.
[0132] Specifically, the vector product of the lines connecting the center point and the four corner points and the sides of the paragraph text boxes connected to them can be expressed as:
[0133] cross_product(P1B,P1K), cross_product(EC,EK), cross_product(P2D,P2K), cross_product(FA,FK), if the above four vector products are all greater than 0, the center point of the paragraph text box corresponding to the paragraph is located inside the target text box.
[0134] All paragraphs whose center points are within the target text box are determined to be paragraphs where the text to be recognized is located. All paragraphs are merged together to obtain the area where the text to be recognized is located that is to be located in this embodiment.
[0135] This embodiment discloses the above-mentioned text positioning method, which first locates the fingertip position, crops the captured image, and then identifies each text line contained in the captured image, determines the paragraph to which each text line belongs, analyzes the area of each paragraph, selects the paragraph text box with the largest area, and then determines the target text box based on the first text box with the largest area and the fingertip position. Each paragraph whose center point is within the target text box is locked as the text content that needs to be identified based on the fingertip locking of the finger in this embodiment. The method provided by this embodiment improves the accuracy of positioning the text to be identified by gradually analyzing the fingertip position information, text line position information and paragraph size information contained in the captured captured image, reduces the requirements for the captured image angle, and provides a guarantee for the smooth recognition of characters.
[0136] The following is a more detailed description of the method in accordance with a specific application example of the method, in conjunction with FIG. 2 to FIG. 7 .
[0137] In step H1, use one or two hands to lock the text to be recognized on the paper text page, as shown in FIG2 .
[0138] Step H2: Use the camera of the smart device to capture an image of the hand locking on the text to be recognized, and transmit the captured image to the text positioning system in the smart device. The positioning system can perform corresponding functions in the form of an application or a mini-program.
[0139] Step H3: The hand detection model stored in the positioning system receives the input captured image, outputs the hand area frame obtained after detecting the hand position in the captured image, crops the hand area frame from the captured image to obtain a hand image, and then performs key point detection on the hand image to obtain the fingertip positions P1 and P2.
[0140] Step H4: The positioning system further identifies the area between the fingertip positions, calculates the midpoint between the two fingertip positions, and calculates the length L of the line connecting the two fingertip positions, and intercepts an area image with a height and width of L, as shown in Figure 3.
[0141] In step H5, the region image is input into the trained text detection model to obtain the text boxes of each text line in the region image. The text lines are then merged through post-processing methods to obtain the text boxes of the paragraphs, as shown in FIG4 .
[0142] Step H6: Find the paragraph text box that intersects the line connecting the upper left fingertip position P1 and the lower right fingertip position P2. For the four sides of the paragraph text box, if any of the four line segments intersects with P1P2, then P1P2 intersects the text box, as shown in Figure 5.
[0143] The specific judgment method is: if line segment AB intersects line segment CD, then the vector product satisfies:
[0144] cross_product(AC,AD)*cross_product(BC,BD)<=0;
[0145] And cross_product(CA,CB)*cross_product(DA,DB)<=0
[0146] Step H7: Select the first text box box_max with the largest area from the paragraph text boxes corresponding to the paragraphs intersecting with P1P2, record the direction of the long side of the text box box_max with the largest area, and obtain a large box target based on P1, P2 and the direction of the long side of the largest text box box_max. It can be obtained that the directions of box_target and box_max are consistent, as shown in Figure 6.
[0147] Step H8: Find all paragraphs whose center points are within box_max. These are the text segments identified by the fingertip (the white box in the figure above) that need to be recognized. The judgment method is: if point K is inside rectangle ABCD, and ABCD is arranged in a clockwise direction, then the cross products cross_product(AB,AK), cross_product(BC,BK), cross_product(CD,CK), and cross_product(DA,DK) are all greater than 0.
[0148] The positioning method provided in this embodiment uses the fingertips of both hands to determine the text area to be recognized. Simply place the fingertip of your left hand at the upper left point of the text line to be recognized, and the fingertip of your right hand at the lower right point of the text line to be recognized. Regardless of the orientation of the text lines in the image, the text area to be recognized can be accurately selected.
[0149] This embodiment discloses a text positioning device, as shown in FIG8 , comprising:
[0150] The region image determination module 710 is used to detect the fingertip position of the fingertip on the text page and determine the region image of the text to be recognized in the text page based on the fingertip position; its function is as described in step S1.
[0151] The text box identification module 720 is used to identify the first text box with the largest area among the paragraph text boxes corresponding to the paragraphs in the regional image; its function is as described in step S2.
[0152] The target text box identification module 730 is used to determine the target text box according to the first text box and the fingertip position; its function is as described in step S3.
[0153] The text area determination module 740 is used to determine each paragraph whose center point is located in the target text box as the paragraph where the text to be recognized is located, and locate the text to be recognized. Its function is as described in step S4.
[0154] Furthermore, the fingertip position includes at least one first fingertip position located at the upper left point of the text page content and at least one second fingertip position located at the lower right point of the text page content; the region image determination module includes: a fingertip position recognition unit, a length calculation unit and a region image determination unit;
[0155] The fingertip position recognition unit is configured to obtain a photographed image of a fingertip on a text page and recognize a first fingertip position and a second fingertip position in the photographed image;
[0156] A length calculation unit, configured to calculate a length value of a midpoint of a connecting line between the first fingertip position and the second fingertip position and the connecting line;
[0157] The area image determination unit is used to use the midpoint of the connecting line as the center point of the cut image and the length of the connecting line as the length and width of the cut image, cut an image from the captured image, and determine the cut image as the area image of the text to be recognized.
[0158] Optionally, the fingertip position recognition unit includes: a shooting subunit, a feature extraction subunit, a cropping subunit and a position detection subunit.
[0159] a photographing subunit, configured to photograph an image of a fingertip positioned on a text page;
[0160] a feature extraction subunit, configured to locate a hand region frame where the hand is located from the captured image;
[0161] a cropping subunit, configured to crop a hand region frame from the captured image to obtain a hand image;
[0162] The position detection subunit is used to perform key point detection on the hand image to obtain a detected first fingertip position and a second fingertip position.
[0163] Furthermore, the fingertip position recognition unit further includes: a correction subunit;
[0164] The correction subunit is configured to detect whether the visual axis of the camera has changed; when a change in the visual axis of the camera is detected, obtain an updated image after the change in the visual axis; extract key points and feature descriptors from the hand image and the updated image respectively; and match the hand image and the updated image using the key points and feature descriptors to obtain a transformation matrix between the hand image and the updated image;
[0165] The first fingertip position in the updated image is calculated using the transformation matrix, and the first fingertip position in the updated image is used as the updated first fingertip position.
[0166] Furthermore, the text box recognition module includes: a text line detection unit, a paragraph merging unit and a screening unit;
[0167] The text line detection unit is configured to input the region image into a trained text detection model to obtain a second text box corresponding to each text line in the region image output by the text detection model;
[0168] The paragraph merging unit is used to identify the paragraphs where each text line is located based on the geometric position relationship between the second text boxes corresponding to each text line, and then obtain the paragraph text boxes corresponding to each paragraph;
[0169] The screening unit is used to screen out the first text box with the largest area from the paragraph text boxes corresponding to each paragraph.
[0170] Optionally, the fingertip positions include a first fingertip position located at an upper left point of the text page content and a second fingertip position located at a lower right point of the text page content;
[0171] The screening unit includes: an intersection segment acquisition subunit and a screening area subunit;
[0172] The intersecting paragraph obtaining subunit is used to obtain a paragraph that intersects with a line connecting the fingertips of the first fingertip position and the second fingertip position;
[0173] The area screening sub-unit is used to screen out the first text box with the largest area from the paragraph text boxes corresponding to the intersecting paragraphs.
[0174] Optionally, the text area determination module includes: a target text box positioning unit, a coordinate acquisition unit, a center point judgment unit and an area determination unit.
[0175] The target text box positioning unit is configured to obtain the long side direction of the first text box; and determine the target text box according to the long side direction of the first text box and the fingertip position;
[0176] The coordinate acquisition unit is used to obtain the coordinates of the four corners of the paragraph text box corresponding to each paragraph in the target text box, and calculate the paragraph center point of the paragraph text box corresponding to each paragraph in turn according to the coordinates of the four corners of the paragraph text box corresponding to each paragraph;
[0177] The center point determination unit is used to sequentially determine whether the center point of the paragraph text box corresponding to each paragraph is located within the target text box;
[0178] The region determination unit is configured to select one or more paragraphs whose center points are located within the target text frame, and determine the text contained in each selected paragraph as the text to be recognized.
[0179] Based on the disclosure of the above-mentioned method and apparatus, this embodiment further discloses an intelligent device, as shown in FIG8 . The intelligent device includes a memory, a processor, and a text positioning program stored in the memory and executable on the processor. When the processor executes the text positioning program, the steps of the text positioning method are implemented.
[0180] FIG9 is a schematic diagram of the structure of the smart device provided in an embodiment of the present application.
[0181] The smart device may include:
[0182] A memory 801 , a processor 802 , and a computer program stored in the memory 801 and executable on the processor 802 .
[0183] When the processor 802 executes the program, the steps of the text positioning method provided in the above embodiment are implemented.
[0184] Furthermore, the smart device further includes:
[0185] The communication interface 803 is used for communication between the memory 801 and the processor 802 .
[0186] The memory 801 is used to store computer programs that can be run on the processor 502 .
[0187] The memory 801 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0188] If the memory 801, processor 802, and communication interface 803 are implemented independently, the communication interface 803, memory 801, and processor 802 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, FIG9 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0189] Optionally, in a specific implementation, if the memory 801, the processor 802 and the communication interface 803 are integrated on a chip, the memory 801, the processor 802 and the communication interface 803 can communicate with each other through an internal interface.
[0190] The processor 802 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0191] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0192] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection having one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0193] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0194] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0195] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0196] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A text positioning method, characterized in that, The method includes: Detecting the fingertip position of a finger on a text page, and determining a region image of the text to be recognized in the text page based on the fingertip position; Identifying a first text box with the largest area in the paragraph text boxes corresponding to each paragraph within the region image; Determining a target text box according to the first text box and the fingertip position; Screening out one or more paragraphs whose paragraph centers are located within the target text box, and determining the text contained in each of the screened paragraphs as the text to be recognized.
2. The text positioning method according to claim 1, characterized in that The fingertip position includes a first fingertip position at the upper left point of the text page content and a second fingertip position at the lower right point of the text page content; The step of detecting the fingertip position of a finger on a text page and determining a region image of the text to be recognized in the text page based on the fingertip position includes: Obtaining a captured image of a finger fingertip on a text page, and identifying the first fingertip position and the second fingertip position in the captured image; Calculating the midpoint of the connection line between the first fingertip position and the second fingertip position and the length value of the connection line; Taking the midpoint of the connection line as the center of the intercepted image, and taking the length value of the connection line as the length and width of the intercepted image, intercepting an image from the captured image, and determining the intercepted image as the region image of the text to be recognized.
3. The text location method according to claim 2, wherein The step of obtaining a captured image of a finger fingertip on a text page and identifying the first fingertip position and the second fingertip position in the captured image includes: Obtaining a captured image of a finger fingertip on a text page, and identifying the first fingertip position and the second fingertip position in one of the captured images; or, respectively identifying the first fingertip position and the second fingertip position in different captured images.
4. The text positioning method according to claim 2, wherein The step of obtaining a captured image of a finger fingertip on a text page and identifying the first fingertip position and the second fingertip position in the captured image includes: Obtaining a captured image of a finger fingertip on a text page, and when identifying the first fingertip position in one captured image, determining other captured images captured within a preset time, so as to identify the second fingertip position according to the captured images captured within the preset time.
5. The text positioning method according to claim 2, characterized in that, The step of obtaining a captured image of a finger fingertip on a text page and identifying the first fingertip position and the second fingertip position in the captured image includes: Obtaining a captured video of a finger fingertip on a text page, and determining a plurality of consecutive frame images from the captured video; Determining a captured image according to the abscissa and ordinate of the fingertip position in the plurality of consecutive frame images, so as to determine the first fingertip position and the second fingertip position in the captured image.
6. The text location method according to claim 5, characterized in that The determining a captured image according to the abscissa and ordinate of the fingertip position in the plurality of consecutive frame images includes: In a preset number of consecutive frame images, if the change ranges of the abscissa and ordinate of the fingertip position in each adjacent frame image are both less than a preset pixel threshold, then among the plurality of consecutive frame images, determining one of the images as the captured image.
7. The text location method according to claim 2, characterized in that The steps of obtaining a captured image of a finger tip on a text page and identifying a first finger tip position and a second finger tip position in the captured image include: Capturing a captured image of a finger tip located on a text page; Locating a hand region box of the position of the hand in the captured image; Cropping the hand region box from the captured image to obtain a hand image; Performing key point detection on the hand image to obtain the detected first finger tip position and second finger tip position.
8. The text positioning method according to claim 7, characterized in that The locating a hand region box of the position of the hand in the captured image includes: Based on a trained hand detection model, determining a hand region box of the position of the hand in the captured image according to the captured image.
9. The text positioning method according to claim 8, characterized in that, The method further includes: Training the hand detection model according to a positive sample set and a negative sample set to obtain a trained hand detection model; wherein, the positive sample set includes pictures of the finger tips of a finger placed on a paper text page and marked with the region box of the position of the hand where the finger tips are located, and the negative sample set includes pictures without a hand.
10. The text location method according to claim 7, wherein After the step of performing key point detection on the hand image to obtain the detected first finger tip position and second finger tip position, it further includes: Detecting whether the optical axis of the camera changes; When it is detected that the optical axis of the camera changes, acquiring an updated image after the optical axis changes; Respectively extracting key points and feature descriptors in the hand image and the updated image; Matching the hand image and the updated image by using the key points and feature descriptors to obtain a transformation matrix between the hand image and the updated image; Calculating the first finger tip position in the updated image by using the transformation matrix, and taking the first finger tip position in the updated image as the updated first finger tip position.
11. The text positioning method according to claim 10, wherein The matching the hand image and the updated image by using the key points and feature descriptors to obtain a transformation matrix between the hand image and the updated image includes: Determining a first image according to the coordinates of the first finger tip position in the hand image and a preset side length; Determining a second image according to the coordinates of the first finger tip position in the updated image and the preset side length; Determining a transformation matrix between the hand image and the updated image according to the key points and feature descriptors in the first image and the second image; Wherein, the length of the preset side length is greater than or equal to one quarter of the length of the text page and less than or equal to the length of the text page; or the length of the preset side length is greater than or equal to one quarter of the width of the text page and less than or equal to the width of the text page.
12. The text location method according to claim 1, characterized in that The step of identifying a first text box with the largest area among the paragraph text boxes corresponding to each paragraph in the region image includes: Inputting the region image into a trained text detection model to obtain second text boxes corresponding to each text line in the region image output by the text detection model; Identifying the paragraphs where each text line is located according to the geometric position relationship between the second text boxes corresponding to each text line, and further obtaining paragraph text boxes corresponding to each paragraph; Selecting a first text box with the largest area from the paragraph text boxes corresponding to each paragraph.
13. The text location method according to claim 12, characterized in that, Identifying the paragraphs where each text line is located based on the geometric positional relationship between the second text boxes corresponding to each text line, and further obtaining the paragraph text boxes corresponding to each paragraph, including: Determining the text lines corresponding to the second text boxes that are adjacent and have no blank lines between them as the same paragraph, so as to obtain the paragraph text boxes corresponding to each paragraph.
14. The text positioning method according to claim 12, characterized in that, Identifying the paragraphs where each text line is located based on the geometric positional relationship between the second text boxes corresponding to each text line, and further obtaining the paragraph text boxes corresponding to each paragraph, including: Determining the position coordinates of the four corner points of the second text box; Determining the paragraph where the text line corresponding to each second text box is located based on at least one of the distance, coincidence degree, and alignment degree between the position coordinates of the four corner points of multiple second text boxes, and further obtaining the paragraph text boxes corresponding to each paragraph.
15. The text positioning method according to claim 12, characterized in that, The fingertip position includes a first fingertip position located at the upper left point of the text page content and a second fingertip position located at the lower right point of the text page content; The step of screening out the first text box with the largest area from the paragraph text boxes corresponding to each paragraph includes: Obtaining the paragraphs that intersect with the connection line of the first fingertip position and the second fingertip position; Screening out the first text box with the largest area from the paragraph text boxes corresponding to the intersecting paragraphs.
16. The text positioning method according to claim 1, wherein The step of determining the target text box based on the first text box and the fingertip position includes: Obtaining the long side direction of the first text box; Determining the target text box based on the long side direction of the first text box and the fingertip position.
17. The text positioning method according to any one of claims 1-16, characterized in that, The step of screening out one or more paragraphs whose paragraph centers are located within the target text box includes: Obtaining the four corner coordinates of the paragraph text boxes corresponding to each paragraph within the target text box; Calculating the paragraph center point of each paragraph text box corresponding to each paragraph in sequence based on the four corner coordinates of the paragraph text boxes corresponding to each paragraph; Judging in sequence whether the paragraph center points of the paragraph text boxes corresponding to each paragraph are located within the target text box, and obtaining one or more paragraphs whose paragraph center points are located within the target text box.
18. A text positioning device, characterized in that, Including: An area image determination module, configured to detect the fingertip position of the finger fingertip on the text page, and determine the area image of the text to be recognized in the text page based on the fingertip position; A text box recognition module, configured to recognize the first text box with the largest area among the paragraph text boxes corresponding to each paragraph within the area image; A target text box recognition module, configured to determine the target text box based on the first text box and the fingertip position; A text area determination module, configured to screen out one or more paragraphs whose paragraph centers are located within the target text box, and determine the text contained in each of the screened paragraphs as the text to be recognized.
19. An intelligent device, characterized in that, Including a memory, a processor, and a text positioning program stored in the memory and executable on the processor. When the processor executes the text positioning program, the steps of the text positioning method according to any one of claims 1-17 are implemented.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the text localization method according to any one of claims 1-17.
Citation Information
Patent Citations
Image document layout recognition method, device and system
CN114170423A
Text content recognition method and device, computer equipment and storage medium
CN115131693A
Finger sentence query method and device, electronic equipment and computer storage medium
CN116740721A
Text recognition method, electronic equipment and storage medium
CN116824588A
Text positioning method and device, equipment and storage medium
CN118015627A
Cited By
Merging misidentified text structures in a document
US20250165703A1