Page positioning method, auxiliary reading method based thereon and application
By combining image features and text information, a page number localization method has been developed, which solves the problem of high page number recognition error rate in existing technologies and achieves high-precision page number localization in complex situations.
Patent Information
- Application Number
- CN202211153522.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-09-21
AI Technical Summary
In existing technologies, page number recognition methods suffer from high error rates and are not robust when faced with a large amount of text information or user-edited information.
By acquiring the image feature vector of the finger-reading image and the page text information, the page number is located in the pre-established image feature table and page data table respectively, and the page number location result is determined by combining the preset conditions. The complementarity of image features and text information is used to improve the location accuracy.
In cases with multiple texts or images, the accuracy and robustness of page number positioning are achieved, improving the reliability of page number recognition.
Smart Images

Figure CN115437504B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of data processing, and more specifically, the embodiments of the present invention relate to a page number positioning method, an auxiliary reading method based thereon, and its application. Background Technology
[0002] This section is intended to provide background or context for embodiments of the invention set forth in the claims. The description herein may include concepts that may be explored, but not necessarily concepts that have been previously conceived or explored. Therefore, unless otherwise stated, what is described in this section is not prior art for the purposes of this application's specification and claims, and is not acknowledged as prior art simply by virtue of its inclusion in this section.
[0003] Throughout the reading process, when readers encounter difficulties with the text, they often need external assistance, such as a dictionary. This disrupts the flow of reading and easily distracts the reader. Assistive reading techniques can effectively address these issues.
[0004] When using assisted reading, after the book is placed, a camera captures an image of the book, and an algorithm analyzes the image to determine which book and page it belongs to. If the reader points to a specific text within the book, the assisted reading system can also identify the text being pointed to and play it aloud to help the reader recognize the words.
[0005] In existing technologies, some page number recognition methods exist. These methods search multiple similar stored pages in a database based on a finger-reading image, then extract feature information from pre-marked areas in both the stored page and the finger-reading image to determine the stored page corresponding to the finger-reading image, and thus the corresponding page number. However, this method relies on image features, and when there is a lot of text or user corrections, it suffers from a high page number recognition error rate and is not robust. Summary of the Invention
[0006] Existing page number recognition methods rely on the similarity of image features for page number location. When faced with a large amount of text information or user corrections, these methods suffer from a high error rate and poor page number recognition performance.
[0007] Therefore, there is a great need for an improved page number positioning method that can accommodate page number positioning needs of various types of content, and can achieve accurate page number positioning in cases with multiple images or multiple texts.
[0008] In this context, embodiments of the present invention aim to provide a page number positioning method, an auxiliary reading method based thereon, and its application.
[0009] In a first aspect of the embodiments of the present application, a page number positioning method is provided, comprising: obtaining a read image; extracting an image feature vector and page text information of the read image; searching in a pre-established image feature table according to the image feature vector to determine a first page number positioning result; searching in a pre-established page data table according to the page text information to determine a second page number positioning result; and obtaining a positioned page number according to the first page number positioning result and the second page number positioning result.
[0010] In an embodiment of the present application, the obtaining of the positioned page number according to the first page number positioning result and the second page number positioning result comprises: if the first page number positioning result and the second page number positioning result are inconsistent, taking the second page number positioning result as the positioned page number when the page text information satisfies a preset condition; and the preset condition comprises that a text quantity in the page text information is greater than or equal to a quantity threshold.
[0011] In an embodiment of the present application, the preset condition further comprises one or more of the following conditions: a similarity score of a text search result obtained based on the page text information is greater than a first score threshold, and a score difference between a text search result corresponding to a highest similarity score and a text search result corresponding to a second highest similarity score is greater than a first difference threshold.
[0012] In an embodiment of the present application, the text in the page text information is printed text.
[0013] In an embodiment of the present application, after the obtaining of the read image, the method further comprises: detecting and removing interference information in the read image to obtain a non-interference read image; and updating the read image with the non-interference read image; and the interference information comprises handwritten text and erasure trace features.
[0014] In an embodiment of the present application, before the searching in the pre-established image feature table according to the image feature vector, the method further comprises: obtaining a standard image of each page of a warehoused book; the standard image is any one of a scanned image and an electronic book image; performing text detection and recognition on each standard image to generate a page data table; wherein the page data of each standard image corresponds to a page number; and performing image feature vector extraction on each standard image to generate an image feature table; wherein the image feature vector of each standard image corresponds to a page number.
[0015] In one embodiment of the present application, the page text information is double-page text information in the read image; the page data table comprises a single-page data table and a double-page data table; correspondingly, the page code positioning method comprises: acquiring the read image; extracting an image feature vector and double-page text information of the read image; searching in a pre-established image feature table according to the image feature vector to determine a first double-page page code positioning result; searching in the double-page data table according to the double-page text information to determine a second double-page page code positioning result; obtaining a positioned double-page page code according to the first double-page page code positioning result and the second double-page page code positioning result; performing page detection on the read image; if double-page page information is obtained through the page detection, a first positioning strategy is executed to determine the positioned page code in the positioned double-page page code; if single-page page information is obtained through the page detection or no page information is obtained through the page detection, a second positioning strategy is executed to determine the positioned page code in the positioned double-page page code.
[0016] In one embodiment of the present application, the read image contains positioning information of a read object fed back by a user; correspondingly, the execution of the first positioning strategy to determine the positioned page code in the positioned double-page page code comprises: determining page information of a page pointed by the user according to the positioning information; and determining the positioned page code in the positioned double-page page code based on the page information of the page pointed by the user.
[0017] In one embodiment of the present application, the page information comprises a page category, a page positioning frame and a page edge key point; the page category comprises a left page and a right page; correspondingly, the execution of the first positioning strategy to determine the positioned page code in the positioned double-page page code comprises: determining a page category of a page pointed by the user according to a relative relationship between the positioning information and the page positioning frame and the page edge key point; and determining the positioned page code in the positioned double-page page code according to the page category of the page pointed by the user.
[0018] In one embodiment of the present application, the read image also contains positioning information of a read object fed back by a user; correspondingly, the execution of the second positioning strategy to determine the positioned page code in the positioned double-page page code comprises: extracting local text information of a region pointed by the user in the read image according to the positioning information; and matching the local text information with page data corresponding to the positioned double-page page code in the single-page data table to determine the positioned page code.
[0019] In a second aspect of the embodiments of the present application, a page code positioning-based auxiliary reading method is provided, comprising: acquiring a pointing image; the pointing image containing positioning information of a pointing object fed back by a user; extracting an image feature vector of the pointing image and page text information; searching in a pre-established image feature table according to the image feature vector to determine a first page code positioning result; searching in a pre-established page data table according to the page text information to determine a second page code positioning result; obtaining a positioning page code according to the first page code positioning result and the second page code positioning result; determining target reading text in the positioning page code according to the positioning information; and performing voice playing on the target reading text.
[0020] In an embodiment of the present application, the determining of the target reading text in the positioning page code according to the positioning information comprises: locating the target reading text in page data corresponding to the positioning page code in the page data table according to the positioning information and / or local text information; wherein the local text information is text information of a user pointing area in the pointing image.
[0021] In an embodiment of the present application, the locating of the target reading text in the page data corresponding to the positioning page code in the page data table according to the positioning information comprises: calculating an affine transformation matrix of the pointing image and a standard image corresponding to the positioning page code; the standard image is used to generate the page data table and the image feature table; converting coordinate information of the positioning information in the standard image according to the positioning information and the affine transformation matrix; and locating the target reading text in the page data corresponding to the positioning page code according to the coordinate information.
[0022] In an embodiment of the present application, the locating of the target reading text in the page data corresponding to the positioning page code in the page data table according to the local text information comprises: searching in the page data corresponding to the positioning page code based on the local text information to obtain a text search result corresponding to a highest similarity score as the target reading text.
[0023] In one embodiment of the present application, locating the target reading text in the page data corresponding to the locating page number in the page data table according to the locating information and the local text information comprises: calculating an affine transformation matrix of the reading image and a standard image corresponding to the locating page number; the standard image is used to generate the page data table and the image feature table; calculating coordinate information of the locating information in the standard image according to the locating information and the affine transformation matrix; locating a first reading text in the page data corresponding to the locating page number according to the coordinate information; searching in the page data corresponding to the locating page number based on the local text information to obtain a second reading text corresponding to a highest similarity score; and determining a target reading text from the first reading text and the second reading text based on the similarity score of the first reading text or the similarity score of the second reading text.
[0024] In one embodiment of the present application, determining the target reading text from the first reading text and the second reading text based on the similarity score of the first reading text or the similarity score of the second reading text comprises: comparing the first reading text with the local text information to obtain the similarity score of the first reading text; if the similarity score of the first reading text is less than a second score threshold, taking the second reading text as the target reading text; or if the similarity score of the second reading text is less than the second score threshold, taking the first reading text as the target reading text; or if searching in the page data corresponding to the locating page number based on the local text information further obtains a third reading text corresponding to a second highest similarity score, taking the first reading text as the target reading text when a similarity score difference between the second reading text and the third reading text is less than a second difference threshold.
[0025] In a third aspect of the embodiments of the present application, a page number locating device is provided, comprising: an imaging device configured to acquire a reading image; an information extraction device configured to extract an image feature vector of the reading image and page text information; a data searching device configured to search in a pre-established image feature table according to the image feature vector to determine a first page number locating result, and search in a pre-established page data table according to the page text information to determine a second page number locating result; and a page number analysis device configured to obtain a locating page number according to the first page number locating result and the second page number locating result.
[0026] In a fourth aspect of the embodiments of the present application, an auxiliary reading device is provided, comprising: an imaging device configured to acquire a finger reading image; the finger reading image containing positioning information fed back by a user through the imaging device; an information extraction device configured to extract an image feature vector and page text information of the finger reading image; a data retrieval device configured to retrieve in a pre-established image feature table according to the image feature vector to determine a first page code positioning result, and retrieve in a pre-established page data table according to the page text information to determine a second page code positioning result; a page code analysis device configured to obtain a positioned page code according to the first page code positioning result and the second page code positioning result; a text positioning device configured to determine target reading text in the positioned page code according to the positioning information; and a speech synthesis device configured to perform speech playing on the target reading text.
[0027] In a fifth aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor; and a memory having stored thereon executable code that, when executed by the processor, causes the processor to perform the method of any one of the preceding aspects.
[0028] In a sixth aspect of the embodiments of the present application, a non-transitory machine readable storage medium having stored thereon executable code that, when executed by a processor of an electronic device, causes the processor to perform the method of any one of the preceding aspects.
[0029] According to the page code positioning method of the embodiments of the present application, the image feature vector and the page text information of the finger reading image can be used as reliable retrieval basis to determine the positioned page code of the finger reading image, so as to improve the accuracy and robustness of the page code positioning. BRIEF DESCRIPTION OF DRAWINGS
[0030] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which a number of embodiments of the present application are illustrated by way of example, in which:
[0031] Figure 1 a block diagram schematically illustrating an exemplary computing system 100 suitable for implementing embodiments of the present application;
[0032] Figure 2A flowchart schematically showing a page number positioning method according to an embodiment of the present application is shown;
[0033] Figure 3 A flowchart schematically showing a page number positioning method according to another embodiment of the present application is shown;
[0034] Figure 4 A flowchart schematically showing an auxiliary reading method according to an embodiment of the present application is shown;
[0035] Figure 5 A flowchart schematically showing a target reading text positioning method according to an embodiment of the present application is shown;
[0036] Figure 6 A flowchart schematically showing a target reading text positioning method according to another embodiment of the present application is shown;
[0037] Figure 7 A structural block diagram schematically showing a page number positioning apparatus according to an embodiment of the present application is shown;
[0038] Figure 8 A structural block diagram schematically showing an auxiliary reading apparatus according to an embodiment of the present application is shown;
[0039] In the drawings, identical or corresponding reference signs indicate identical or corresponding parts. DETAILED DESCRIPTION
[0040] The principles and spirits of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only so that those skilled in the art can better understand and implement the present application, and do not limit the scope of the present application in any way. On the contrary, these embodiments are provided so that the present disclosure is more thorough and complete, and the scope of the present disclosure is fully conveyed to those skilled in the art.
[0041] Figure 1 A block diagram of an exemplary computing system 100 suitable for implementing embodiments of the present application is shown. As Figure 1As shown, the computing system 100 can include a central processing unit (CPU) 101, a random access memory (RAM) 102, a read only memory (ROM) 103, a system bus 104, a hard disk controller 105, a keyboard controller 106, a serial interface controller 107, a parallel interface controller 108, a display controller 109, a hard disk 110, a keyboard 111, a serial peripheral 112, a parallel peripheral 113, and a display 114. Of these devices, the CPU 101, RAM 102, ROM 103, hard disk controller 105, keyboard controller 106, serial interface controller 107, parallel interface controller 108, and display controller 109 are coupled to the system bus 104. The hard disk 110 is coupled to the hard disk controller 105, the keyboard 111 is coupled to the keyboard controller 106, the serial peripheral 112 is coupled to the serial interface controller 107, the parallel peripheral 113 is coupled to the parallel interface controller 108, and the display 114 is coupled to the display controller 109. It will be appreciated that Figure 1 The structural diagram described is merely for the purpose of example, and is not a limitation on the scope of the present application. In some cases, some devices can be added or reduced according to specific circumstances.
[0042] Those skilled in the art know that the embodiments of the present application can be implemented as a system, a method or a computer program product. Therefore, the present disclosure can be embodied in the form of entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to herein as "circuitry", "module" or "system". In addition, in some embodiments, the present application can also be embodied in the form of a computer program product in one or more computer readable media, which contains computer readable program code.
[0043] Any combination of one or more computer readable medium can be employed. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer readable storage medium can include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.
[0044] Computer readable signal media can include a propagated data signal with computer readable program code embodied therein. For example, a propagated signal can be an electromagnetic signal, an optical signal, and / or any other suitable type of signal. Such a propagated signal can carry computer readable program code in the form of electrical signals, optical signals, and / or magnetic
[0045] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0046] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). These network connections are
[0047] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0048] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0049] These computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0050] According to an embodiment of the present application, a page number positioning method and device are provided.
[0051] In this document, it is to be understood that the number of elements in the figures is for illustration only and does not limit the scope of the application. Any of the named elements can be duplicated or omitted.
[0052] The principles and spirit of the present application will be explained in detail below with reference to several representative embodiments of the present application. SUMMARY
[0054] The present inventors have found that the auxiliary reading technology needs to rely on a page number recognition method. However, the page number recognition method in the prior art relies on the similarity between the image captured by the camera and the image features to position the corresponding page number. When there is a lot of text information or user modification traces, the page number cannot be accurately recognized.
[0055] Although the image features of the page with a lot of text information cannot support the accuracy of page number recognition, such a page provides a large amount of rich text information, which can be used as supplementary information to make up for the deficiency of image features, and thus improve the breadth of the basis used for page number positioning. Similarly, for the page with poor text information and rich image features, the image features can assist the text information to achieve high-precision page number positioning.
[0056] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be described in detail below.
[0057] Overview of Application Scenarios
[0058] The page number positioning method of the embodiments of the present application is suitable for various reading scenarios or products, including but not limited to a point-and-read machine. In addition, the page number positioning method of the embodiments of the present application can be applied to the reading scenario of a paper book, and can also be applied to the reading scenario of an electronic book.
[0059] Exemplary Method
[0060] The principles and spirit of the present application will be explained in detail below with reference to several representative embodiments of the present application. Figure 2A method of page number positioning according to an exemplary embodiment of the present application will be described. It should be noted that the above application scenarios are merely shown for the convenience of understanding the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0061] In the technical solutions of the present application, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0062] The page number positioning method provided by the present application can comprise:
[0063] In step 201, a reading image is acquired. The reading image contains book content. The book here can be a paper book or an electronic book, without limitation.
[0064] In some embodiments, the reading image can be an image collected by an imaging device or a frame image extracted from a video in a reading scene. The collection of the reading image can be triggered by the user's action, for example, when a specific object exists in the shooting picture, the imaging device is triggered to collect the reading image. The specific object can be a finger tip or a reading pen.
[0065] In step 202, the image feature vector and the page text information of the reading image are extracted.
[0066] In the embodiments of the present application, the image features of the reading image include but are not limited to color features, and further include color block contour features.
[0067] In some embodiments, the text of the page text information is printed text. The printed text includes but is not limited to the following fonts: Songti, Kaishu, Lishu, Heiti and Xingkai, etc. In actual application, the handwritten text in the reading image is often human-added text information, and the pre-established page data table contains the original page text information of each page. Therefore, the detected handwritten text may interfere with the text retrieval result.
[0068] Based on the above consideration, in some embodiments, after acquiring the reading image, the following steps can be performed to eliminate the interference information in the reading image, specifically comprising:
[0069] Detecting and removing the interference information in the reading image to obtain a non-interference reading image;
[0070] Updating the reading image with the non-interference reading image; wherein the interference information includes handwritten text and erasure trace features.
[0071] In the above embodiment, the removing of the interference information can be performed before step 202, and the image processing technique is used to process the read image to obtain a read image without interference, and then the image feature vector and the page text information are extracted.
[0072] In another embodiment, the removing of the interference information can also be performed synchronously with step 202, for example, when the page text information is extracted, the extracted text is judged by font, if it is a handwritten text, it is deleted, and if it is a printed text, it is retained.
[0073] It can be understood that the removing of the interference information is only an example provided in the embodiment of the present application, and does not constitute the only limitation of the present application.
[0074] In step 203, the image feature vector is searched in a pre-established image feature table to determine the first page code positioning result.
[0075] In step 204, the page text information is searched in a pre-established page data table to determine the second page code positioning result.
[0076] Before steps 203 and 204 are performed, the books in the warehouse need to be processed in advance to establish the required page data table and image feature table.
[0077] Specifically, before the image feature vector is searched in the pre-established image feature table, the following steps can be included: obtaining a standard image of each page of the book in the warehouse; the standard image is any one of a scanned image and an electronic book image; performing text detection and recognition on each standard image to generate a page data table; wherein the page data of each standard image corresponds to the page code; performing image feature vector extraction on each standard image to generate an image feature table; wherein the image feature vector of each standard image corresponds to the page code.
[0078] It should be noted that the page code positioning method of the present application is not only suitable for paper books, but also for electronic books. When the book in the warehouse is a paper book, the standard image of each page of the book in the warehouse can be obtained by scanning the book page, and the image feature table and the page data table are generated based on the scanned image. When the book in the warehouse is an electronic book, each page of the book in the warehouse is already saved in electronic format, and it can be directly imported to obtain the standard image of each page.
[0079] It should be noted that the embodiment of the present application does not have strict requirements on the execution timing of steps 203 and 204, and in actual application, step 204 can be executed before step 203 or in parallel with step 203, which is not limited herein.
[0080] In step 205, the positioning page code is obtained according to the first page code positioning result and the second page code positioning result.
[0081] In actual application, if the first page number positioning result and the second page number positioning result are consistent, it indicates that the page number positioning results are consistent based on the image features or the page text information, and any one of the first page number positioning result and the second page number positioning result can be taken as the final positioning page number.
[0082] Of course, in actual application, there are cases where the text retrieval result and the image retrieval result are inconsistent. If the first page number positioning result and the second page number positioning result are inconsistent, the final positioning page number can be determined according to a preset condition. For example, when the page text information meets the preset condition, the second page number positioning result is taken as the positioning page number. The preset condition includes that the number of texts in the page text information is greater than or equal to a number threshold.
[0083] When the number of texts in the page text information is greater than the number threshold, it indicates that the content in the page is mainly text information, that is, the page text information provides a large amount of rich information that can be used as the basis for page number positioning. Therefore, it is more reliable to take the second page number positioning result based on the page text information as the final positioning page number.
[0084] It should be noted that the number threshold can be a preset value, and in actual application, the value of the number threshold can also be adjusted according to actual conditions.
[0085] In some embodiments, the preset condition can also include one or more of the following conditions: a similarity score of a text retrieval result retrieved based on the page text information is greater than a first score threshold, and a score difference between a text retrieval result corresponding to a highest similarity score and a text retrieval result corresponding to a second highest similarity score is greater than a first difference threshold.
[0086] When the score difference between the text retrieval result corresponding to the highest similarity score and the text retrieval result corresponding to the second highest similarity score is less than or equal to the first difference threshold, it indicates that there are two text retrieval results with similar page text information but different page numbers. In this case, the possibility that the text retrieval result corresponding to the second highest similarity score becomes the actual page number positioning result is greatly increased. If the second page number positioning result based on the page text information is used at this time, the error rate is high.
[0087] In actual application, the first score threshold and the first difference threshold can be adjusted according to actual conditions, which is not limited here.
[0088] The page number positioning method can search the corresponding page number positioning result in the image feature table and the page data table according to the corresponding feature type based on the image feature vector and the page text information of the image, and then analyze the page number positioning results identified based on the two feature types to determine the positioning page number corresponding to the image. Whether the image content is mainly picture information or text information, the present application can ensure reliable search basis, i.e., the image feature vector and the page text information of the image, to determine the positioning page number of the image, thereby improving the accuracy and robustness of the page number positioning.
[0089] The page number positioning method provided by another embodiment of the present application will be described below. Figure 3 The page number positioning method provided by another embodiment of the present application will be described below.
[0090] Figure 3 The flowchart of the page number positioning method according to another embodiment of the present application is schematically shown, and the page number positioning method in this embodiment can include the following steps. Figure 3
[0091] In step 301, the image is acquired.
[0092] In this embodiment, the content of step 301 is consistent with that of step 201 in the above embodiment, and will not be described here again.
[0093] In step 302, the image feature vector and the double-page text information of the image are extracted.
[0094] In this embodiment, the in-storage book is designed as double pages, so that the imaging device can collect the double-page content when the user reads, and the double-page text information can be extracted based on the image. Compared with the single-page text information, the text quantity is increased, and accordingly, the text search result is more accurate.
[0095] Correspondingly, in step 302, the image feature vector of the image can also be the image feature vector corresponding to the double pages.
[0096] In step 303, the first double-page page number positioning result is determined by searching the image feature table according to the image feature vector.
[0097] In step 304, the second double-page page number positioning result is determined by searching the double-page data table according to the double-page text information.
[0098] Correspondingly, in this embodiment, the image feature table, the single-page data table and the double-page data table can also be established in advance before step 303 is performed.
[0099] The double-page data table can be obtained by connecting the page data of two single pages located on the same opening page in the single-page data table to form the double-page data table after the single-page data table is established.
[0100] It should be noted that the execution timing of steps 303 and 304 is not strictly required in the embodiment of the application, and in actual application, step 304 can be performed before step 303 or both steps are performed in parallel, which is not uniquely limited herein.
[0101] Since the double-page text information is used for retrieval, step 304 obtains the double-page page code, and in the subsequent steps, the single-page page code needs to be further located as the locating page code.
[0102] In step 305, the locating double-page page code is obtained according to the first double-page page code locating result and the second double-page page code locating result.
[0103] The specific implementation of step 305 in this embodiment can refer to the implementation of step 205 in the above embodiment, which is not described herein.
[0104] In step 306, the page detection is performed on the image, and the corresponding locating strategy is performed according to the page detection result to determine the locating page code in the locating double-page page code.
[0105] In step 306, the trained target detection model can be used to detect the page information of the page, and the page information can include one or more of the following information: page category, page locating frame and page edge key point.
[0106] In this embodiment, the warehouse book is designed as a double-opening page, and the page category can include a left page and a right page.
[0107] In other embodiments, the warehouse book can also be designed as a three-opening page, and in this case, the page category can also include a left page, a middle page and a right page.
[0108] When the page detection result only contains single-page information, or even fails to detect the page information, it indicates that the page is not completely in the collection image of the imaging device, or the page tilt angle is too large, resulting in page detection failure. At this time, it is difficult to simply rely on the positioning information in the finger reading image to distinguish the left page and the right page, and thus it is difficult to determine the positioning page number.
[0109] When the page detection result contains double-page information, it indicates that the finger reading image contains complete page features. Then, the positioning information in the finger reading image is used to determine the page category pointed by the user, and the positioning page number is determined.
[0110] Based on the above two page detection results, different positioning strategies can be used to determine the positioning page number in the positioning double-page page number, as follows:
[0111] If the page detection obtains double-page information, a first positioning strategy is executed to determine the positioning page number in the positioning double-page page number.
[0112] If the page detection obtains single-page information or fails to obtain page information, a second positioning strategy is executed to determine the positioning page number in the positioning double-page page number.
[0113] The positioning strategies used for different page detection results are described as follows.
[0114] The positioning information of the finger reading object fed back by the user in the finger reading image obtained in step 301 can be the position information of the position of the fingertip of the user's finger, or the position information selected by the point reading pen.
[0115] The first positioning strategy corresponding to successful detection (the page detection obtains double-page information) is as follows:
[0116] The page information of the page pointed by the user is determined according to the positioning information.
[0117] The positioning page number in the positioning double-page page number is determined based on the page information of the page pointed by the user.
[0118] The page information includes: page category, page positioning frame, and page edge key point; the page category includes: left page and right page.
[0119] Correspondingly, the specific execution mode of the first positioning strategy can include:
[0120] The page category of the page pointed by the user is determined according to the relative relationship between the positioning information and the page positioning frame and the page edge key point.
[0121] The positioning page number in the positioning double-page page number is determined according to the page category of the page pointed by the user.
[0122] In the above case, the target detection model can detect the page positioning frame and the page edge key point of the left page and the right page, that is, through the page information of the double pages, the left page position and the right page position in the reading image can be clearly distinguished. On this basis, by determining whether the positioning information is located in the page positioning frame of which side page in the double pages, it can be determined whether the user points to the left page or the right page, and then the positioning page number is known.
[0123] The second positioning strategy corresponding to the detection failure (the page detection obtains the single page information or fails to detect the page information) is as follows:
[0124] According to the positioning information, local text information of a user pointing area in the reading image is extracted;
[0125] The local text information is matched with the page data corresponding to the positioning double page number in the single page data table, so as to determine the positioning page number.
[0126] In the above case, the complete left page and the right page cannot be detected, and therefore it is difficult to determine the page category pointed by the user according to the positioning information. At this time, the local text information of the user pointing area in the reading image is extracted, wherein the size of the pointing area can be a preset area size. According to the local text information, text retrieval is performed in the single page data table, so as to determine the positioning page number.
[0127] Further, the range of the text retrieval can be limited in the positioning double page number determined in step 305. By narrowing the range of the text retrieval, the efficiency of the page number positioning can be improved, and the interference of the page text information of other page numbers on the text retrieval is excluded, so as to improve the accuracy of the page number positioning.
[0128] Based on the page number positioning method of any one of the foregoing embodiments, one embodiment of the present application further provides an auxiliary reading method for playing the text pointed by the user in voice according to the positioning information of the reading object fed back by the user, so as to help the user recognize the text.
[0129] The auxiliary reading method provided in the embodiment will be described below in combination with Figure 4 The specific implementation of the auxiliary reading method will be described.
[0130] Please refer to Figure 4 The auxiliary reading method provided in the embodiment can include:
[0131] In step 401, a reading image is obtained. The reading image contains book page content, and can also contain positioning information of a reading object fed back by a user. The book page herein can be a paper book page or an electronic book page, and is not limited to be unique.
[0132] In the embodiment, the pointing reading object can be text pointed by a finger of a user, or text specified by a user through a pointing pen or the like.
[0133] The content of step 401 in the embodiment is consistent with that of step 201 in the foregoing embodiment, and will not be described here again.
[0134] In step 402, an image feature vector of the pointing reading image and page text information are extracted.
[0135] The content of step 402 in the embodiment is consistent with that of step 202 in the foregoing embodiment, and will not be described here again.
[0136] In step 403, the image feature vector is searched in a pre-established image feature table to determine a first page code positioning result.
[0137] The content of step 403 in the embodiment is consistent with that of step 203 in the foregoing embodiment, and will not be described here again.
[0138] In step 404, the page text information is searched in a pre-established page data table to determine a second page code positioning result.
[0139] The content of step 404 in the embodiment is consistent with that of step 204 in the foregoing embodiment, and will not be described here again.
[0140] In step 405, the first page code positioning result and the second page code positioning result are used to obtain a positioned page code.
[0141] The content of step 405 in the embodiment is consistent with that of step 205 in the foregoing embodiment, and will not be described here again.
[0142] In step 406, target reading text in the positioned page code is determined according to the positioning information.
[0143] Since the pointing reading image can be interfered by light, paper deformation, distance, erasing traces and the like, the result of directly detecting and recognizing the text of the pointing reading image has a high error rate, therefore, in the embodiment, the target reading text is determined in the single page data table according to the positioning information, and is used for voice playing, so as to ensure the accuracy of the played text.
[0144] Specifically, the target reading text can be located in the page data corresponding to the positioned page code in the page data table according to the positioning information and / or local text information.
[0145] The local text information is text information of a region in the pointing image that is pointed by the user, for example, after the user points to a position on a page by using a finger or a pointing pen as an indicating tool, text information of a region in the pointing image that is circled with the position as a base point.
[0146] In some embodiments, the circled size can be set as a fixed size to assist in determining the local text information. In other embodiments, the circled text can be further segmented, and text information of a text segment closest to the positioning information is taken as the local text information.
[0147] In step 407, the target reading text is played by voice.
[0148] By using the auxiliary reading method of the above embodiments, the positioning page number that the user points to a page can be located according to the pointing image, and the target reading text can be located in the page data of the positioning page number according to the positioning information of the pointing object that is fed back by the user. Since the target reading text is text information recorded in a single page data table, compared with directly detecting and recognizing the text of the pointing image, the text information is more accurate.
[0149] The method of locating the target reading text according to the positioning information alone is described below.
[0150] Referring to Figure 5 , the method of locating the target reading text according to the positioning information includes:
[0151] In step 501, an affine transformation matrix of the pointing image and a standard image corresponding to the positioning page number is calculated. The standard image is used to generate a page data table and an image feature table.
[0152] When the warehoused book is a paper book, the standard image can be a scanning image of each page in the paper book. When the warehoused book is an electronic book, the standard image can be an electronic book image corresponding to each page in the electronic book.
[0153] Due to different angles of the imaging device, the standard image is an image collected by directly facing the page under ideal light conditions, while the pointing image can be regarded as a deformed image of the standard image, that is, the two images can be converted by an affine transformation matrix.
[0154] In step 502, coordinate information of the positioning information in the standard image is calculated according to the positioning information and the affine transformation matrix.
[0155] Due to the differences between the standard image and the reading image, the coordinate information of the page position pointed to by the positioning information in the standard image deviates from the positioning information. Therefore, the coordinate information of the positioning information corresponding to the reading image in the standard image can be calculated by the above affine transformation matrix, thereby determining the target reading text corresponding to the area pointed to by the user.
[0156] In step 503, the target reading text is located in the page data corresponding to the location page number based on the coordinate information.
[0157] The target text localization method provided in this embodiment converts the localization information into coordinate information in the corresponding standard image through an affine transformation matrix. The target text is then determined based on the text information recorded in the standard image, effectively correcting the distortion problem of the finger-reading image caused by the angle of the imaging device, thereby improving the accuracy of target text localization.
[0158] The following explains the method of locating the target text based solely on local text information.
[0159] The method for locating target text based on local text information is as follows:
[0160] Based on local text information, a search is performed on the page data corresponding to the located page number, and the text search result with the highest similarity score is used as the target reading text.
[0161] Although there is a certain percentage of text errors when directly performing text detection and recognition on the finger-reading image, in practical applications, text retrieval can be performed on the correct page data based on this local text information to retrieve the text information that best matches the local text information in the page data corresponding to the location page number, and use it as the target text location.
[0162] Because the scope of text retrieval is limited to the page data corresponding to the located page number, the search range is narrowed and a large amount of interfering information is eliminated, providing a certain degree of error tolerance. Even if there is a certain proportion of text errors in the local text information, reliable target reading text can still be found.
[0163] An embodiment of the present invention also provides another method for locating target reading text by combining location information and local text information.
[0164] The following is combined Figure 6 The method for locating the target text is explained.
[0165] In step 601, the affine transformation matrix between the pointing image and the standard image corresponding to the page number is calculated. The standard image is used to generate the page data table and the image feature table.
[0166] In this embodiment, the content of step 601 is consistent with that of step 501 in the previous embodiment, which will not be repeated here.
[0167] In step 602, the coordinate information of the positioning information in the standard image is calculated according to the positioning information and the affine transformation matrix.
[0168] In this embodiment, the content of step 602 is consistent with that of step 502 in the previous embodiment, which will not be repeated here.
[0169] In step 603, the first reading text is located in the page data corresponding to the positioning page number according to the coordinate information.
[0170] In this embodiment, the content of step 603 can refer to the content of step 503 in the previous embodiment, which will not be repeated here.
[0171] In step 604, the second reading text corresponding to the highest similarity score is obtained by searching the page data corresponding to the positioning page number based on the local text information.
[0172] In this embodiment, based on the local text information, multiple text search results can be obtained, and the obtained multiple text search results are arranged in order from high to low similarity score. The first text search result in the order is taken as the second reading text, the second text search result in the order is taken as the third reading text, and so on.
[0173] In step 605, the target reading text is determined in the first reading text and the second reading text based on the similarity score of the first reading text or the similarity score of the second reading text.
[0174] Specifically, the execution process of step 605 includes but is not limited to the following three cases:
[0175] In the first case, the target reading text is determined based on the similarity score of the first reading text. Then, the similarity score of the first reading text can be obtained by comparing the first reading text with the local text information. If the similarity score of the first reading text is less than the second score threshold, the second reading text is taken as the target reading text.
[0176] In the first case, the similarity score of the first reading text is less than the second score threshold, which means that the difference between the first reading text determined by the affine transformation matrix and the local text information is too large, which indicates that the coordinate information has a large deviation, resulting in a high possibility of error in the located target reading text. In this case, the error rate of taking the first reading text as the target reading text is high, so the second reading text is determined as the target reading text.
[0177] In the second case, the target reading text is determined based on the similarity score of the second reading text. If the similarity score of the second reading text is less than the second score threshold, the first reading text is taken as the target reading text.
[0178] Similarly to the first case, when the similarity score of the second reading text is less than the second score threshold, it indicates that the second reading text is too different from the local text information, which means that the error rate of taking the second reading text as the target reading text is high in this case. Therefore, the first reading text is determined as the target reading text.
[0179] In the third case, the target reading text is determined based on the similarity scores of the second reading text and the third reading text. Then, the page data corresponding to the positioning page number is searched based on the local text information to obtain the third reading text corresponding to the second highest similarity score. When the difference between the similarity scores of the second reading text and the third reading text is less than the second difference threshold, the first reading text is taken as the target reading text.
[0180] In the third case, when the difference between the similarity scores of the second reading text and the third reading text is small, it indicates that the third reading text and the second reading text have similar likelihood of becoming the page number positioning result that conforms to the actual situation. Therefore, the likelihood of the third reading text being the correct target reading text is greatly increased, and it is difficult to distinguish which one of the two is the correct target reading text. At this time, the error rate of taking the second reading text as the target reading text is high.
[0181] It should be noted that the above is only an exemplary description of several possible execution cases of step 605. In actual application, there are other execution manners, for example:
[0182] In the fourth case, before the coordinate information is converted by using the affine transformation matrix, the feature points of the standard image and the reading image can be calculated respectively by using the ORB (Oriented FAST and Rotated BRIEF) feature detection algorithm, and feature matching is performed. The affine transformation matrix is calculated according to the feature matching result. Further, high-quality feature points can be selected to perform the action of feature matching, wherein the high-quality feature points refer to the feature points with a feature distance within a range of 2 times the minimum distance. In step 605, the proportion of the number of feature points participating in the calculation of the affine transformation matrix to the total number of detected feature points can be combined to determine whether to take the first reading text or the second reading text as the target reading text.
[0183] Exemplarily, when the proportion of the number of feature points participating in the affine transformation matrix calculation to the total number of detected feature points is less than a preset proportion threshold, it is indicated that the accuracy of the affine transformation matrix is low, the accuracy of the coordinate information converted based on the affine transformation matrix is low, and thus the possibility of the first reading text being the target reading text is low, and the second reading text is taken as the target reading text.
[0184] In actual application, step 605 further includes a fifth case, in which the first reading text is consistent with the second reading text, and either of the two reading texts can be taken as the target reading text.
[0185] The embodiment provides a method for positioning a target reading text by combining positioning information and local text information, which positions the reading text by using two types of information (positioning information and text information), analyzes two positioning results, and selects a positioning result that is more consistent with an actual situation as the target reading text by competing with each other, thereby improving the accuracy and reliability of the target reading text.
[0186] Exemplary Device
[0187] After introducing the method of the exemplary embodiment of the application, next, reference is made to Figure 7 The page code positioning device of the exemplary embodiment of the application is introduced.
[0188] The page code positioning device 700 provided by the embodiment of the application comprises:
[0189] An imaging device 701 is configured to acquire a finger reading image.
[0190] An information extraction device 702 is connected to the imaging device 701 and configured to extract an image feature vector and page text information of the finger reading image.
[0191] A data retrieval device 703 is connected to the information extraction device 702 and configured to retrieve in a pre-established image feature table according to the image feature vector to determine a first page code positioning result, and retrieve in a pre-established page data table according to the page text information to determine a second page code positioning result.
[0192] A page code analysis device 704 is connected to the data retrieval device 703 and configured to obtain a positioned page code according to the first page code positioning result and the second page code positioning result.
[0193] Further, the page code positioning device 700 can further comprise a storage device (not shown in the figure) connected to the page code analysis device 704 and storing the image feature table and the page information table.
[0194] Further, the page number analysis device 704 can be further configured to: if the first page number positioning result and the second page number positioning result are inconsistent, take the second page number positioning result as the positioning page number when the page text information meets a preset condition; wherein the preset condition comprises that the number of texts in the page text information is greater than or equal to a number threshold.
[0195] The page number positioning device provided by the embodiment can ensure reliable search basis, i.e., the image feature vector of the read image and the page text information, to determine the positioning page number of the read image, thereby improving the accuracy and robustness of page number positioning.
[0196] Next, reference is made to Figure 8 An auxiliary reading device according to an exemplary embodiment of the present application is introduced.
[0197] The auxiliary reading device 800 provided by the embodiment comprises:
[0198] An imaging device 801 is configured to acquire a read image; the read image contains positioning information fed back by a user through the imaging device;
[0199] An information extraction device 802 is connected with the imaging device 801 and configured to extract an image feature vector and page text information of the read image;
[0200] A data search device 803 is connected with the information extraction device 802 and configured to search in a pre-established image feature table according to the image feature vector to determine a first page number positioning result, and search in a pre-established page data table according to the page text information to determine a second page number positioning result;
[0201] A page number analysis device 804 is connected with the data search device 803 and configured to obtain a positioning page number according to the first page number positioning result and the second page number positioning result;
[0202] A text positioning device 805 is connected with the page number analysis device 804, the imaging device 801 and the information extraction device 802 respectively and configured to determine target reading texts in the positioning page number according to the positioning information;
[0203] A speech synthesis device 806 is connected with the text positioning device 805 and configured to perform speech playing on the target reading texts.
[0204] Further, the auxiliary reading device 800 can further comprise a storage device (not shown in the figure) which stores an image feature table and a page information table.
[0205] Corresponding to any of the method embodiments described above, the embodiments of the present application also provide an electronic device, comprising: a processor; and a memory having stored thereon executable code which, when executed by the processor, causes the processor to perform the method according to any of the embodiments described above.
[0206] Alternatively, the present application can also be implemented as a non-transitory machine readable storage medium (or computer readable storage medium, or machine readable storage medium) having stored thereon executable code (or computer program, or computer instruction code), which, when executed by a processor of an electronic device (or electronic device, server, etc.), causes the processor to perform part or all of the steps of the method according to the above.
[0207] It should be noted that although in the above detailed description, several devices or sub-devices corresponding to the methods are mentioned, such division is merely not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more devices described above can be embodied in one device. Conversely, the features and functions of one device described above can be further divided into devices embodied by multiple devices.
[0208] In addition, although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all of the shown operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step execution, and / or one step can be divided into multiple steps.
[0209] The use of the verb "comprise", "comprising", and words of similar meaning in the application file does not exclude the presence of elements other than those listed in the application file. The article "a" or "an" preceding an element does not exclude the presence of multiple such elements.
[0210] Although the spirit and principles of the present application have been described with reference to several specific embodiments, it should be understood that the present application is not limited to the disclosed specific embodiments, and the division of aspects does not mean that the features in these aspects cannot be combined for the benefit. This division is only for the convenience of expression. The present application is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the appended claims. The scope of the appended claims is the broadest interpretation, so as to include all such modifications and equivalent structures and functions.
Claims
1. A page number positioning method, characterized by, The method comprises: obtaining a finger reading image; extracting an image feature vector and page text information of the finger reading image; retrieving in a pre-established image feature table according to the image feature vector to determine a first page code positioning result; retrieving in a pre-established page data table according to the page text information to determine a second page code positioning result; obtaining a positioned page code according to the first page code positioning result and the second page code positioning result.
2. The page number positioning method of claim 1, wherein, The method of obtaining a positioned page code according to the first page code positioning result and the second page code positioning result comprises: if the first page code positioning result and the second page code positioning result are inconsistent, then when the page text information meets a preset condition, taking the second page code positioning result as the positioned page code. The preset condition comprises that the number of texts in the page text information is greater than or equal to a number threshold.
3. The method according to claim 2, wherein the preset condition further comprises one or more of the following conditions: a similarity score of a text retrieval result obtained based on the page text information is greater than a first score threshold, and a score difference between a text retrieval result corresponding to a highest similarity score and a text retrieval result corresponding to a second highest similarity score is greater than a first difference threshold.
4. The method according to claim 1, wherein the texts in the page text information are printed texts. After obtaining the finger reading image, the method further comprises: detecting and removing interference information in the finger reading image to obtain a non-interference finger reading image; 5. The page number positioning method of claim 4, wherein, updating the finger reading image with the non-interference finger reading image. The interference information comprises handwritten texts and erasure trace features. Before retrieving in the pre-established image feature table according to the image feature vector, the method further comprises: obtaining a standard image of each page of a warehoused book; the standard image is any one of a scanned image and an electronic book image; 6. The page number positioning method of claim 1, wherein, performing text detection and recognition on each standard image to generate a page data table; wherein the page data of each standard image corresponds to a page code; performing image feature vector extraction on each standard image to generate an image feature table; wherein the image feature vector of each standard image corresponds to a page code.
7. The method according to claim 1, wherein the page text information is double-page text information in the finger reading image; the page data table comprises a single-page data table and a double-page data table; correspondingly, the method comprises: obtaining a finger reading image; extracting an image feature vector and double-page text information of the finger reading image; retrieving in a pre-established image feature table according to the image feature vector to determine a first double-page page code positioning result; retrieving in the double-page data table according to the double-page text information to determine a second double-page page code positioning result; obtaining a positioned double-page page code according to the first double-page page code positioning result and the second double-page page code positioning result; performing page detection on the finger reading image; if the page detection obtains double-page page information, then executing a first positioning strategy to determine the positioned page code in the positioned double-page page code. If the page detection obtains the page information of the single page or the page detection does not obtain the page information, a second positioning strategy is executed to determine the positioning page number in the positioning double-page page number.
8. The page number positioning method of claim 7, wherein the positioning information of the pointing object fed back by the user is contained in the pointing image. Accordingly, the execution of the first positioning strategy to determine the positioning page number in the positioning double-page page number comprises: determining the page information of the page pointed by the user according to the positioning information; and determining the positioning page number in the positioning double-page page number based on the page information of the page pointed by the user.
9. The page number positioning method of claim 8, wherein the page information comprises a page category, a page positioning frame and a page edge key point, and the page category comprises a left page and a right page. Accordingly, the execution of the first positioning strategy to determine the positioning page number in the positioning double-page page number comprises: determining the page category of the page pointed by the user according to the relative relationship between the positioning information and the page positioning frame and the page edge key point; and determining the positioning page number in the positioning double-page page number according to the page category of the page pointed by the user.
10. The page number positioning method of claim 7, wherein the pointing image further contains the positioning information of the pointing object fed back by the user. Accordingly, the execution of the second positioning strategy to determine the positioning page number in the positioning double-page page number comprises: extracting local text information of a region pointed by the user in the pointing image according to the positioning information; and matching the local text information with page data corresponding to the positioning double-page page number in the single-page data table to determine the positioning page number. The method comprises: obtaining a pointing image; the pointing image contains the positioning information of the pointing object fed back by the user; 11. An auxiliary reading method based on page number positioning, characterized by, extracting an image feature vector and page text information of the pointing image; searching in a pre-established image feature table according to the image feature vector to determine a first page number positioning result; searching in a pre-established page data table according to the page text information to determine a second page number positioning result; obtaining a positioning page number according to the first page number positioning result and the second page number positioning result; determining target reading text in the positioning page number according to the positioning information; playing the target reading text by voice. The determination of the target reading text in the positioning page number according to the positioning information comprises: locating the target reading text in page data corresponding to the positioning page number in the page data table according to the positioning information and / or local text information; the local text information is text information of a region pointed by the user in the pointing image.
12. The page number based positioning aided reading method according to claim 11, wherein, The locating of the target reading text in the page data corresponding to the positioning page number in the page data table according to the positioning information and the local text information comprises: calculating an affine transformation matrix of the pointing image and a standard image corresponding to the positioning page number; the standard image is used to generate the page data table and the image feature table. 13. The page number based positioning aided reading method according to claim 12, characterized in that, According to the positioning information and the affine transformation matrix, coordinate information of the positioning information in the standard image is calculated; According to the coordinate information, first reading text in the page data corresponding to the positioning page code is located; Based on the local text information, second reading text corresponding to the highest similarity score is obtained by searching in the page data corresponding to the positioning page code; Based on the similarity score of the first reading text or the similarity score of the second reading text, the target reading text is determined in the first reading text and the second reading text.
14. The page number based positioning aided reading method according to claim 12, wherein, According to the local text information, the target reading text is located in the page data corresponding to the positioning page code in the page data table, including: Based on the local text information, the text search result corresponding to the highest similarity score is obtained by searching in the page data corresponding to the positioning page code, as the target reading text.
15. The page number based positioning aided reading method according to claim 12, wherein, According to the positioning information and the local text information, the target reading text is located in the page data corresponding to the positioning page code in the page data table, including: The affine transformation matrix of the reading image and the standard image corresponding to the positioning page code is calculated; the standard image is used to generate the page data table and the image feature table; According to the positioning information and the affine transformation matrix, coordinate information of the positioning information in the standard image is calculated; According to the coordinate information, first reading text in the page data corresponding to the positioning page code is located; Based on the local text information, second reading text corresponding to the highest similarity score is obtained by searching in the page data corresponding to the positioning page code; Based on the similarity score of the first reading text or the similarity score of the second reading text, the target reading text is determined in the first reading text and the second reading text.
16. The page number based positioning aided reading method according to claim 15, wherein, The target reading text is determined in the first reading text and the second reading text based on the similarity score of the first reading text or the similarity score of the second reading text, including: The first reading text is compared with the local text information to obtain the similarity score of the first reading text; if the similarity score of the first reading text is less than a second score threshold, the second reading text is taken as the target reading text; Or If the similarity score of the second reading text is less than a second score threshold, the first reading text is taken as the target reading text; Or If based on the local text information, searching in the page data corresponding to the positioning page code also obtains third reading text corresponding to the second highest similarity score, when the similarity score difference between the second reading text and the third reading text is less than a second difference threshold, the first reading text is taken as the target reading text.
17. A page number positioning apparatus, characterized by, Including: An imaging device for acquiring a reading image; An information extraction device for extracting an image feature vector and page text information of the reading image; An information extraction device for extracting an image feature vector and page text information of the reading image; The data retrieval device is configured to search in a pre-established image feature table according to the image feature vector to determine a first page number positioning result, and search in a pre-established page data table according to the page text information to determine a second page number positioning result. The page number analysis device is configured to obtain the positioned page number according to the first page number positioning result and the second page number positioning result.
18. An auxiliary reading device, characterized in that The method comprises: An imaging device is configured to acquire a finger reading image; The finger reading image contains positioning information fed back by a user through the imaging device; An information extraction device is configured to extract an image feature vector and page text information of the finger reading image; The data retrieval device is configured to search in a pre-established image feature table according to the image feature vector to determine a first page number positioning result, and search in a pre-established page data table according to the page text information to determine a second page number positioning result; The page number analysis device is configured to obtain the positioned page number according to the first page number positioning result and the second page number positioning result. A text positioning device is configured to determine target reading text in the positioned page number according to the positioning information; A speech synthesis device is configured to perform speech playing on the target reading text.
19. An electronic device, comprising: The method comprises: A processor; And A memory having executable code stored thereon, when the executable code is executed by the processor, the processor executes the method as claimed in any one of claims 1-16.
20. A non-transitory machine readable storage medium having executable code stored thereon, when the executable code is executed by a processor of an electronic device, the processor executes the method as claimed in any one of claims 1-16.
Citation Information
Patent Citations
Page number recognition method, device, reading robot and computer readable storage medium
CN110532964A
Paper book page number recognition method and device, family education machine and storage medium
CN110647648A