A method and system for implementing a user fingertip positioning problem
By digitizing teaching aids and using fingertip coordinates to assist in mapping similar pages, the problem of low positioning accuracy in photo searches is solved, and the precise positioning of questions is achieved in natural page photography scenes.
Patent Information
- Application Number
- CN202210510201.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-11
AI Technical Summary
The existing technology is affected by text layout and photo integrity during the process of taking and searching for photos, resulting in a decrease in the accuracy of the question positioning and the original questions required by users cannot be accurately given.
By scanning the paper teaching aids into electronic versions, selecting the area of each question to form the coordinates of the question blocks, and depositing the entire page of OCR text and question blocks into the digital library of the teaching aids according to the page. When a user takes a photo, he uses the similar structure of the fingertip coordinate position near the row to assist in the mapping of the fingertip coordinates, thereby achieving accurate positioning of the topic in the natural page photo scene.
In the natural page photo scene, accurately positioning the questions on the user's fingertips is improved, and the accuracy of the positioning of the questions is less affected by the quality of the photo taken.
Smart Images

Figure CN117115848B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software development, and particularly to a method and system for implementing user fingertip positioning of questions. Background Art
[0002] With the popularization of intelligent devices such as mobile phones, tablets, intelligent table lamps, learning machines, etc. and the development of educational software, various application scenarios such as taking pictures to search for questions and fingertip querying of questions have emerged.
[0003] Currently, the question positioning in the industry is basically to search according to the question-cutting area after layout analysis. This method is affected by the text layout and the integrity of the photo, such as the tilting, deformation, and missing pages of the picture caused by non-standard photographing; the missed cutting and wrong cutting of the question-cutting by layout analysis. The above situations will all lead to a reduction in the accuracy of question positioning and cannot accurately give the original question required by the user. Based on this, there is an urgent need for a method and system for implementing user fingertip positioning of questions to improve the accuracy of question positioning. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for implementing user fingertip positioning of questions to achieve accurate user fingertip question positioning.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A method for implementing user fingertip positioning of questions includes:
[0007] Obtaining a plurality of first text line blocks, first line block coordinates, and question block coordinates in an electronic teaching aid book; the first text line block is a block obtained by enclosing several adjacent characters in each line of the electronic teaching aid book with a rectangle, the first line block coordinates are the four vertex coordinates of the rectangle, and the question block coordinates are the vertex coordinates of a block obtained by enclosing a single question in the electronic teaching aid book with a rectangle;
[0008] Storing the first text line block, the first line block coordinates, and the question block coordinates page by page to obtain a digital library of the teaching aid book, denoted as the library;
[0009] Taking a picture to obtain a picture containing a fingertip and a question to be searched, denoted as the photographed picture, establishing a second coordinate system with the upper left corner of the photographed picture as the origin, and obtaining the coordinates of the fingertip point in the photographed picture; the fingertip in the photographed picture is positioned to the question to be searched;
[0010] Performing OCR recognition on the photographed picture to obtain a plurality of second text line blocks in the photographed picture, determining the coordinates of the second text line blocks according to the second coordinate system, denoted as second line block coordinates, and the second text line block is a block obtained by enclosing several adjacent characters in each line of the photographed picture with a rectangle;
[0011] According to the positions of the texts in the photographed image, splice the multiple second text line blocks to obtain the OCR text of the whole page; according to the OCR text of the whole page, retrieve the pages similar to the OCR text of the whole page from the library to form a set of similar pages;
[0012] Perform standardization processing on both the photographed image after OCR recognition and the set of similar pages to obtain the processed photographed image and the processed set of similar pages. Obtain multiple text lines and line coordinates according to the processed photographed image; the text line refers to a line composed of text line blocks in each line of the processed photographed image, and the line coordinate is the coordinate of the text line;
[0013] Determine the line located by the fingertip point according to the line coordinate and the fingertip point coordinate, denoted as the fingertip line;
[0014] Perform text similarity comparison between the fingertip line and each text line of each page in the processed set of similar pages, and take the text line with the highest similarity in the comparison result of each page, denoted as the similar line;
[0015] Merge the text of the fingertip line with the text of the upper and lower lines of the fingertip line to obtain the first text; merge the text of each similar line with the text of the upper and lower lines of the similar line to obtain multiple merged second texts; perform similarity comparison between each second text and the first text, and take the page where the first text with the highest comparison result is located, denoted as the book page;
[0016] Determine the question where the similar line is located according to the coordinate of the similar line and the coordinates of multiple question blocks in the book page.
[0017] The present invention also provides a system for realizing the positioning of questions by the user's fingertip, including:
[0018] A first text line block acquisition module, configured to acquire multiple first text line blocks, first line block coordinates, and question block coordinates in an electronic teaching aid book; the first text line block is a block obtained by enclosing several adjacent texts in each line of the electronic teaching aid book with a rectangle, the first line block coordinate is the four vertex coordinates of the rectangle, and the question block coordinate is the vertex coordinate of a block obtained by enclosing a single question in the electronic teaching aid book with a rectangle;
[0019] A library storage module, configured to store the first text line block, the first line block coordinate, and the question block coordinate page by page to obtain a digital library of the teaching aid book, denoted as the library;
[0020] The fingertip point coordinate acquisition module is used to take a photo to obtain a picture containing the fingertip and the question to be searched, denoted as the photo graph. A second coordinate system is established with the upper left corner of the photo graph as the origin, and the coordinates of the fingertip point in the photo graph are obtained; the fingertip in the photo graph is positioned to the question to be searched;
[0021] The second text acquisition module is used to perform OCR recognition on the photo graph to obtain multiple second text line blocks in the photo graph, and determine the coordinates of the second text line blocks according to the second coordinate system, denoted as the second line block coordinates. The second text line block is a block obtained by enclosing several adjacent characters in each row of the photo graph with a rectangle;
[0022] The similar page set acquisition module is used to splice multiple second text line blocks according to the positions of the characters in the photo graph to obtain the full-page OCR text; according to the full-page OCR text, retrieve the pages similar to the full-page OCR text from the library to form a similar page set;
[0023] The standardization processing module is used to perform standardization processing on both the photo graph after OCR recognition and the similar page set to obtain the processed photo graph and the processed similar page set; obtain multiple text lines and line coordinates according to the processed photo graph; the text line refers to a line composed of text line blocks in each row of the processed photo graph, and the line coordinate is the coordinate of the text line;
[0024] The fingertip line acquisition module is used to determine the line where the fingertip point is located according to the line coordinates and the fingertip point coordinates, denoted as the fingertip line;
[0025] The similar line acquisition module is used to compare the text similarity between the fingertip line and each text line of each page in the processed similar page set, and select the text line with the highest similarity in the comparison result of each page, denoted as the similar line;
[0026] The page acquisition module is used to merge the text of the fingertip line with the text of the upper and lower lines of the fingertip line to obtain the first text; merge the text of each similar line with the text of the upper and lower lines of the similar line to obtain multiple merged second texts; compare the similarity between each second text and the first text, and select the page where the first text with the highest comparison result is located, denoted as the page;
[0027] The question determination module is used to determine the question where the similar line is located according to the coordinates of the similar line and the coordinates of multiple question blocks in the page.
[0028] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0029] The present invention provides a method and system for realizing the positioning of questions by a user's fingertip. First, a paper-based teaching auxiliary book is scanned into an electronic version, and the areas of each question in the teaching auxiliary book are framed to form question block coordinates. The OCR text of the entire page and the question blocks are stored in the digital library of the teaching auxiliary book by page. On the basis of the digitization of the teaching auxiliary book, the OCR content of the user's photographed image is used to search for a set of similar pages in the library. By using the similar structures of the lines adjacent to the fingertip coordinate position, the mapping of similar pages of the fingertip row coordinates is assisted, so as to realize the accurate positioning of questions in the natural page photographing scenario. Compared with the traditional method, the present application can realize the accurate positioning of the user's fingertip questions, and is less affected by the quality of the photographed image. It only needs to capture the text OCR of the lines adjacent to the fingertip coordinate position in the photographed image for comparison. Brief Description of the Drawings
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 It is a flowchart of the method for realizing the positioning of questions by a user's fingertip provided in Embodiment 1 of the present invention;
[0032] Figure 2 It is the photographed image provided in Embodiment 1 of the present invention;
[0033] Figure 3 It is the photographed image after OCR recognition provided in Embodiment 1 of the present invention;
[0034] Figure 4 It is the photographed image after standardization processing provided in Embodiment 1 of the present invention;
[0035] Figure 5 It is the similar page provided in Embodiment 1 of the present invention;
[0036] Figure 6 It is the similar page provided in Embodiment 1 of the present invention;
[0037] Figure 7 It is the similar page containing similar similar lines provided in Embodiment 1 of the present invention;
[0038] Figure 8 It is the similar page containing similar similar lines provided in Embodiment 1 of the present invention;
[0039] Figure 9 It is the photographed image that circles the context of the fingertip row provided in Embodiment 1 of the present invention;
[0040] Figure 10The photographed image that demarcates the similar line context provided by Embodiment 1 of the present invention;
[0041] Figure 11 The photographed image that demarcates the similar line context provided by Embodiment 1 of the present invention. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] The purpose of the present invention is to provide a method and system for realizing the positioning of a user's fingertip to a question, so as to realize accurate positioning of the user's fingertip to the question.
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0045] Embodiment 1
[0046] This embodiment provides a method for realizing the positioning of a user's fingertip to a question, including:
[0047] S1. Obtain a plurality of first text line blocks, first line block coordinates, and question block coordinates in an electronic teaching aid book;
[0048] Wherein, the first text line block is a block obtained by enclosing several adjacent characters in each line of the electronic teaching aid book with a rectangle, the first line block coordinates are the four vertex coordinates of the rectangle, and the question block coordinates are the four vertex coordinates of a block obtained by enclosing a single question in the electronic teaching aid book with a rectangle;
[0049] Specifically, step S1 may include the following steps:
[0050] S11. Scan the teaching aid book into an electronic version, and perform OCR recognition on the electronic teaching aid book page by page to obtain a first teaching aid book, and the first teaching aid book contains a plurality of first text line blocks;
[0051] S12. Establish a first coordinate system with the upper left corner of the first teaching aid book as the origin, and obtain the line block coordinates according to the first coordinate system;
[0052] S13. Use a rectangle to frame out the coordinate area of a single question from the first teaching aid book to obtain a question block area, and determine the question block coordinates of the question block area according to the first coordinate system.
[0053] S2. Store the first text line block, the first line block coordinates, and the question block coordinates page by page to obtain a digital library of teaching reference books, denoted as the library;
[0054] The above steps S1 and S2 are the process of digitizing teaching reference books. Next, a fingertip question positioning scheme is introduced, including steps S3 - S10.
[0055] S3. Take a photo to obtain a picture containing the fingertip and the question to be searched, denoted as the photo. Establish a second coordinate system with the upper left corner of the photo as the origin, and obtain the coordinates of the fingertip point in the photo; the fingertip in the photo locates the question to be searched;
[0056] Please refer to Figure 2 , in this embodiment, the fingertip point coordinates are (X: 630, Y: 1085).
[0057] S4. Perform OCR recognition on the photo to obtain multiple second text line blocks in the photo. Determine the coordinates of the second text line blocks according to the second coordinate system, denoted as the second line block coordinates. The second text line block is a block obtained by enclosing several adjacent characters in each row of the photo with a rectangle. For details, please refer to Figure 3 ;
[0058] S5. Piece together multiple second text line blocks according to the positions of the characters in the photo to obtain the OCR text for the entire page; retrieve pages similar to the OCR text for the entire page from the library according to the OCR text for the entire page, and form a set of similar pages;
[0059] After obtaining the set of similar pages, the first text line block, the first line block coordinates, and the question block coordinates in the similar pages can be obtained. As shown in Figure 5 and Figure 6 shown, Figure 5 and Figure 6 are both similar pages;
[0060] S6. Perform standardization processing on the photo after OCR recognition and the set of similar pages to obtain the processed photo and the processed set of similar pages. Obtain multiple text lines and line coordinates according to the processed photo; the text line refers to a line composed of text line blocks in each row of the processed photo, and the line coordinates are the coordinates of the text line;
[0061] Affected by the user's photo - taking angle, the text after OCR recognition may be fragmented or skewed. As shown in Figure 3 shown, so it is necessary to perform rotation processing on the OCR recognition result to ensure that the text is horizontal, and then stretch the fragmented text of the OCR coordinates according to the character spacing (empirical threshold) of the printed text to obtain the complete OCR line coordinates and OCR text lines. The result is as shown in Figure 4It should be noted that in this example, the method of standardizing the similar pages in the set of similar pages is the same as the method of standardizing the photographed images. Therefore, the standardization process of similar pages will not be elaborated here.
[0062] S7. Determine the row where the fingertip point is located according to the row coordinate and the fingertip point coordinate, denoted as the fingertip row;
[0063] Specifically, step S7 may include:
[0064] Judge whether the fingertip point is within the row coordinate area according to the fingertip point coordinate;
[0065] If so, determine the row where the fingertip point is located as the fingertip row;
[0066] If not, determine the row that is located at the upper left of the fingertip point and is closest to the fingertip point as the fingertip row. For example, Figure 4 , the text in the row where the fingertip coordinate is located is "When x exceeds 100, for every 5 yuan increase in the daily rent of each vehicle, the number of rented sightseeing vehicles will decrease by 1. Given that all sightseeing vehicles are available every day".
[0067] S8. Compare the fingertip row with each text row of each page in the processed set of similar pages, and take the text row with the highest similarity in the comparison result of each page, denoted as the similar row;
[0068] Traverse the text rows in the processed set of similar pages in step S6 and the fingertip row in step S7 from top to bottom for text similarity comparison, and take the row with the highest similarity in the comparison result as the similar row. They are respectively Figure 7 "When x exceeds 100, for every 5 yuan increase in the daily rent of each vehicle, rent" in Figure 8 "When x exceeds 100, for every 5 yuan increase in the daily rent of each vehicle, the number of rented sightseeing vehicles will decrease by 1. Given that all sightseeing vehicles are available every day" in
[0069] Among them, the method used for text similarity matching is the text similarity algorithm, that is, by comparing the single-line text structure through the text similarity algorithm, the most similar structure on each page in the set of similar pages is found.
[0070] S9. Merge the text of the fingertip row with the text of the upper and lower rows of the fingertip row to obtain the first text; merge the text of each similar row with the text of the upper and lower rows of the similar row to obtain multiple merged second texts; compare the similarity of each second text with the first text, and take the page where the first text with the highest comparison result is located, denoted as the book page;
[0071] Among them, the text of the upper and lower rows of the fingertip row can be referred to Figure 9, see the text above and below the similar lines Figure 10 and Figure 11 . Figure 10 The similarity between the text in and the first text is 0.68, Figure 11 The similarity between the text in and the first text is 0.97. Therefore, Figure 11 is the page corresponding to the photographed image.
[0072] S10. Determine the question where the similar line is located according to the coordinates of the similar line and the coordinates of multiple question blocks in the page.
[0073] In the above steps, the page in the book library corresponding to the user's photographed image has been determined. Traverse the question block areas in the page, and the question block area where the similar line in the page in step S8 is located is the determined question. For example, Figure 6 As can be seen, the coordinate area where the similar line "When x exceeds 100, for every 5 yuan increase in the daily rent of each vehicle, the number of rented sightseeing vehicles will decrease by 1. Given that all sightseeing vehicles are available every day" is located is serial number 4, and the original question is serial number 4.
[0074] It should be noted that in this embodiment, the text structure determined by line structure matching maps the fingertip line in the photographed image to the line in the search result page, so as to determine the location of the question.
[0075] The prerequisite for the present invention is the digitization of teaching aids. After scanning the paper teaching aids into electronic versions, the area of each question is framed to form question block coordinates, and the OCR text and question blocks of the entire page are stored in the teaching aid digitization book library according to pages. On the basis of the digitization of teaching aids, the OCR content of the user's photographed image is used to search for a set of similar pages in the book library, and the similar structure of the lines adjacent to the fingertip coordinate position is used to assist the mapping of the similar pages of the fingertip line coordinates, so as to achieve accurate positioning of the questions in the natural page photographing scenario.
[0076] Compared with the traditional method, this method can achieve accurate positioning of the user's fingertip questions, and is less affected by the quality of the photographed image. It only needs to capture the OCR of the text of the lines adjacent to the fingertip coordinate position in the photographed image for comparison.
[0077] Embodiment 2
[0078] This embodiment provides a system for realizing user fingertip positioning of questions, including:
[0079] The first text line block acquisition module M1 is used to acquire multiple first text line blocks, first line block coordinates, and question block coordinates in the electronic teaching supplementary book; the first text line block is a block obtained by enclosing several adjacent characters in each line of the electronic teaching supplementary book with a rectangle, the first line block coordinates are the four vertex coordinates of the rectangle, and the question block coordinates are the vertex coordinates of a block obtained by enclosing a single question in the electronic teaching supplementary book with a rectangle;
[0080] The library storage module M2 is used to store the first text line blocks, the first line block coordinates, and the question block coordinates page by page to obtain a digital library of the teaching supplementary book, denoted as the library;
[0081] The fingertip point coordinate acquisition module M3 is used to take a photo of a picture containing a fingertip and a question to be searched, denoted as the photo graph. A second coordinate system is established with the upper left corner of the photo graph as the origin, and the coordinates of the fingertip point in the photo graph are acquired; the fingertip in the photo graph is positioned to the question to be searched;
[0082] The second text acquisition module M4 is used to perform OCR recognition on the photo graph to obtain multiple second text line blocks in the photo graph, and determine the coordinates of the second text line blocks according to the second coordinate system, denoted as the second line block coordinates. The second text line block is a block obtained by enclosing several adjacent characters in each line of the photo graph with a rectangle;
[0083] The similar page set acquisition module M5 is used to splice multiple second text line blocks according to the positions of the characters in the photo graph to obtain the full-page OCR text; according to the full-page OCR text, retrieve the pages similar to the full-page OCR text from the library to form a similar page set;
[0084] The standardization processing module M6 is used to perform standardization processing on the photo graph after OCR recognition to obtain the processed photo graph; obtain multiple text lines and line coordinates according to the processed photo graph; the text line refers to a line composed of text line blocks in each line of the processed photo graph, and the line coordinates are the coordinates of the text line;
[0085] The fingertip line acquisition module M7 is used to determine the line located by the fingertip point according to the line coordinates and the fingertip point coordinates, denoted as the fingertip line;
[0086] The similar line acquisition module M8 is used to compare the text similarity between the fingertip line and each text line of each page in the similar page set, and select the text line with the highest similarity in the comparison results of each page, denoted as the similar line;
[0087] A page acquisition module M9 is used to merge the text of the fingertip line with the text of the upper and lower lines of the fingertip line to obtain a first text; merge the text of each of the similar lines with the text of the upper and lower lines of the similar line to obtain a plurality of merged second texts; compare the similarity of each of the second texts with the first text, and take the page where the first text with the highest comparison result is located as the page.
[0088] A question determination module M10 is used to determine the question where the similar line is located according to the coordinates of the similar line and the coordinates of a plurality of question blocks in the page.
[0089] Optionally, the first text line block acquisition module specifically includes:
[0090] A first teaching assistant book scans the teaching assistant book into an electronic version, and performs OCR recognition on the electronic version of the teaching assistant book page by page to obtain a first teaching assistant book, which contains a plurality of text line blocks.
[0091] A line block coordinate acquisition sub-module is used to establish a first coordinate system with the upper left corner of the first teaching assistant book as the origin, and obtain the line block coordinates according to the first coordinate system.
[0092] A question block coordinate acquisition sub-module is used to frame the coordinate area of a single question from the first teaching assistant book with a rectangle to obtain a question block area, and determine the question block coordinates of the question block area according to the first coordinate system.
[0093] Optionally, the standardization processing module specifically includes:
[0094] An OCR photographed image acquisition sub-module is used to record the photographed image after OCR recognition as an OCR photographed image.
[0095] A rotated photographed image acquisition sub-module is used to rotate the OCR photographed image so that the text in the OCR photographed image is horizontal to obtain a rotated photographed image.
[0096] A processed photographed image acquisition sub-module is used to stretch the fragmented text line blocks in each row of the rotated photographed image according to the character spacing in the rotated photographed image to obtain a processed photographed image.
[0097] Optionally, the fingertip line acquisition module specifically includes:
[0098] A judgment sub-module is used to judge whether the fingertip point is located in the line coordinate area according to the fingertip point coordinates.
[0099] If so, determine the line where the fingertip point is located as the fingertip line.
[0100] If not, determine the fingertip row that is located at the upper left of the fingertip point and is the closest to the fingertip point.
[0101] For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, refer to the description in the method section.
[0102] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for implementing a user fingertip positioning question, characterized in that, Including: Obtaining a plurality of first text line blocks, first line block coordinates, and question block coordinates in an electronic teaching aid book; the first text line block is a block obtained by enclosing several adjacent characters in each line of the electronic teaching aid book with a rectangle, the first line block coordinates are the four vertex coordinates of the rectangle, and the question block coordinates are the vertex coordinates of a block obtained by enclosing a single question in the electronic teaching aid book with a rectangle; Storing the first text line blocks, the first line block coordinates, and the question block coordinates page by page to obtain a digital library of the teaching aid book, denoted as the library; Taking a photo to obtain a picture containing a fingertip and a question to be searched, denoted as the photo picture, establishing a second coordinate system with the upper left corner of the photo picture as the origin, and obtaining the coordinates of the fingertip point in the photo picture; the fingertip in the photo picture is positioned to the question to be searched; Performing OCR recognition on the photo picture to obtain a plurality of second text line blocks in the photo picture, and determining the coordinates of the second text line blocks according to the second coordinate system, denoted as the second line block coordinates, where the second text line block is a block obtained by enclosing several adjacent characters in each line of the photo picture with a rectangle; Splicing the plurality of second text line blocks according to the positions of the characters in the photo picture to obtain the full-page OCR text; retrieving pages similar to the full-page OCR text from the library according to the full-page OCR text to form a set of similar pages; Performing standardization processing on both the photo picture after OCR recognition and the set of similar pages to obtain the processed photo picture and the processed set of similar pages, and obtaining a plurality of text lines and line coordinates according to the processed photo picture; the text line refers to a line composed of text line blocks in each line of the processed photo picture, and the line coordinates are the coordinates of the text line; Determining the line located by the fingertip point according to the line coordinates and the fingertip point coordinates, denoted as the fingertip line; Performing text similarity comparison between the fingertip line and each text line of each page in the processed set of similar pages, and taking the text line with the highest similarity in the comparison results of each page, denoted as the similar line; Merging the text of the fingertip line with the text of the upper and lower lines of the fingertip line to obtain a first text; Merging the text of each similar line with the text of the upper and lower lines of the similar line to obtain a plurality of merged second texts; Performing similarity comparison between each second text and the first text, and taking the page where the first text with the highest comparison result is located, denoted as the book page; Determining the question where the similar line is located according to the coordinates of the similar line and the plurality of question block coordinates in the book page.
2. The method according to claim 1, characterized in that, The obtaining of the plurality of text line blocks, line block coordinates, and question block coordinates in the electronic teaching aid book specifically includes: Scanning the teaching aid book into an electronic version, and performing OCR recognition on the electronic version of the teaching aid book page by page to obtain a first teaching aid book, where the first teaching aid book contains a plurality of the first text line blocks; Establishing a first coordinate system with the upper left corner of the first teaching aid book as the origin, and obtaining the line block coordinates according to the first coordinate system; Use a rectangle to frame the coordinate area of a single question from the first supplementary teaching material to obtain a question block area, and determine the question block coordinates of the question block area according to the first coordinate system.
3. The method according to claim 1, characterized in that Perform standardization processing on the photographed image after OCR recognition to obtain the processed photographed image, specifically including: Record the photographed image after OCR recognition as the OCR photographed image; Rotate the OCR photographed image to keep the text in the OCR photographed image horizontal to obtain the rotated photographed image; According to the character spacing in the rotated photographed image, stretch the fragmented text line blocks in each row of the rotated photographed image to obtain the processed photographed image.
4. The method according to claim 1, characterized in that, The specific method for determining the line where the fingertip is located according to the line coordinates and the fingertip point coordinates includes: Judge whether the fingertip is located within the line coordinate area according to the fingertip point coordinates; If so, determine the line where the fingertip is located as the fingertip line; If not, determine the line located at the upper left of the fingertip and closest to the fingertip as the fingertip line.
5. The method according to claim 1, wherein Use the text similarity algorithm to perform the text similarity comparison.
6. A system for implementing a user fingertip positioning question, characterized in that, Including: A first text line block acquisition module, configured to acquire a plurality of first text line blocks, first line block coordinates, and question block coordinates in the electronic supplementary teaching material; the first text line block is a block obtained by using a rectangle to enclose several adjacent characters in each row of the electronic supplementary teaching material, the first line block coordinates are the four vertex coordinates of the rectangle, and the question block coordinates are the vertex coordinates of the block obtained by using a rectangle to enclose a single question in the electronic supplementary teaching material; A library storage module, configured to store the first text line blocks, the first line block coordinates, and the question block coordinates page by page to obtain a digital library of the supplementary teaching material, denoted as the library; A fingertip point coordinate acquisition module, configured to take a photo of a picture including the fingertip and the question to be searched, denoted as the photographed image, establish a second coordinate system with the upper left corner of the photographed image as the origin, and acquire the coordinates of the fingertip point in the photographed image; the fingertip in the photographed image is positioned to the question to be searched; A second text acquisition module, configured to perform OCR recognition on the photographed image to obtain a plurality of second text line blocks in the photographed image, and determine the coordinates of the second text line blocks according to the second coordinate system, denoted as the second line block coordinates, where the second text line block is a block obtained by using a rectangle to enclose several adjacent characters in each row of the photographed image; A similar page set acquisition module, configured to splice a plurality of the second text line blocks according to the positions of the text in the photographed image to obtain the OCR text of the whole page; retrieve the pages similar to the OCR text of the whole page from the library according to the OCR text of the whole page to form a similar page set; A standardization processing module, configured to perform standardization processing on both the photographed image after OCR recognition and the similar page set to obtain the processed photographed image and the processed similar page set; obtain a plurality of text lines and line coordinates according to the processed photographed image; the text line refers to a line composed of the text line blocks in each row of the processed photographed image, and the line coordinates are the coordinates of the text line; A fingertip row acquisition module, configured to determine the row where the fingertip point is located according to the row coordinate and the fingertip point coordinate, denoted as the fingertip row; A similar row acquisition module, configured to perform a text similarity comparison between the fingertip row and each text row of each page in the processed similar page set, and take the text row with the highest similarity in the comparison result of each page, denoted as the similar row; A page acquisition module, configured to merge the text of the fingertip row with the text of the upper and lower rows of the fingertip row to obtain a first text; Merge the text of each similar row with the text of the upper and lower rows of the similar row to obtain multiple merged second texts; Perform a similarity comparison between each second text and the first text, and take the page where the first text with the highest comparison result is located, denoted as the page; A question determination module, configured to determine the question where the similar row is located according to the coordinate of the similar row and the coordinates of multiple question blocks in the page; 7. The system according to claim 6, wherein The first text row block acquisition module specifically includes: The first teaching assistant book scans the teaching assistant book into an electronic version and performs OCR recognition on the electronic version of the teaching assistant book page by page to obtain the first teaching assistant book, and the first teaching assistant book contains multiple text row blocks; A row block coordinate acquisition sub-module, configured to establish a first coordinate system with the upper left corner of the first teaching assistant book as the origin, and obtain the row block coordinates according to the first coordinate system; A question block coordinate acquisition sub-module, configured to use a rectangle to frame the coordinate area of a single question from the first teaching assistant book to obtain a question block area, and determine the question block coordinates of the question block area according to the first coordinate system.
8. The system according to claim 6, wherein The standardization processing module specifically includes: An OCR photographed image acquisition sub-module, configured to denote the photographed image after OCR recognition as the OCR photographed image; A rotated photographed image acquisition sub-module, configured to rotate the OCR photographed image to keep the text in the OCR photographed image horizontal to obtain a rotated photographed image; A processed photographed image acquisition sub-module, configured to stretch the fragmented text row blocks in each row of the rotated photographed image according to the character spacing in the rotated photographed image to obtain a processed photographed image.
9. The system according to claim 6, wherein The fingertip row acquisition module specifically includes: A judgment sub-module, configured to judge whether the fingertip point is located within the row coordinate area according to the fingertip point coordinate; If so, determine the row where the fingertip point is located as the fingertip row; If not, determine the row located at the upper left of the fingertip point and closest to the fingertip point as the fingertip row.
10. The system according to claim 6, characterized in that, The text similarity comparison is performed using a text similarity algorithm.
Citation Information
Patent Citations
Test question searching method and test question searching apparatus applied to electronic terminal
CN105975561A
Whole-page question searching method and terminal
CN111666474A