Text acquisition method, query method, device, storage medium and program product

By building a resource library and applying fuzzy matching rules in the dictionary pen to determine the text to be queried, the problem of inaccurate image and text recognition results in the dictionary pen is solved, and higher quality information query is achieved.

CN120804292APending Publication Date: 2025-10-17HEFEI IFLYTEK TOYCLOUD TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510967937.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the prior art, when a dictionary pen scans text images, the recognition results may contain errors or irrelevant information, resulting in poor accuracy of the query results and failing to meet user needs.

Method used

Build a resource library, store the first text in each resource as the image-text recognition result or modified text to be optimized, and the second text as the set text to be queried. Use fuzzy matching rules to search for the initial text or its modified text in the resource library, determine the text to be queried, and perform knowledge query based on the text to be queried.

Benefits of technology

The quality of information query in image and text recognition scenarios is improved, ensuring the accuracy and reliability of query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804292A_ABST
    Figure CN120804292A_ABST
Patent Text Reader

Abstract

The invention discloses a text acquisition method, a query method, a device, a storage medium and a program product, and relates to the technical field of information processing, the method comprises the following steps: pre-constructing a resource library, a first text in each resource in the resource library being a to-be-optimized image-text recognition result or a modified text of the image-text recognition result, and a second text in each resource in the resource library being a modified text of the to-be-optimized image-text recognition result; the second text in each resource is a to-be-queried text corresponding to the first text in the resource; the query result obtained by performing knowledge query based on the to-be-queried text is superior to the query result obtained by performing knowledge query based on the first text, and based on the resource library, the text recognition result is obtained by performing text recognition on the scanned text image, and then the text recognition result is used as the initial text; the resource library is searched for the resource containing the initial text or the modified text of the initial text, and the second text in at least part of the searched at least one resource is determined as the to-be-queried text corresponding to the text image, so that the information query quality in the image-text recognition scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and in particular to a text acquisition method, a query method, a device, equipment, a storage medium and a program product. BACKGROUND

[0002] With the development of artificial intelligence, some auxiliary learning electronic products have emerged, such as a dictionary pen. A student can obtain corresponding query results by scanning a text area on a medium (a book, a plane, etc.) using a dictionary pen. The dictionary pen captures a text picture on the medium through a camera at a pen tip, performs text recognition on the text picture through optical character recognition (OCR), obtains a text to be queried, and then performs a query on the text to be queried to obtain a query result.

[0003] However, due to various conditions in the scanned text picture, the recognition result can have errors or irrelevant information, and directly using the text recognition result as the text to be queried can result in poor quality of the query result, which cannot meet the query requirements of the user. Therefore, how to improve the accuracy of the query result has become a technical problem to be solved. SUMMARY

[0004] In view of the above problems, the present application provides a text acquisition method, a query method, a device, equipment, a storage medium and a program product to improve the information query quality in a graphic text recognition scenario. The specific solutions are as follows:

[0005] The first aspect of the present application provides a text acquisition method, comprising:

[0006] performing text recognition on the scanned text image to obtain an initial text;

[0007] finding a target text in a resource library, the target text being the initial text or a modified text obtained by modifying a partial text of the initial text; the resource library storing a plurality of resources, a first text in each resource being a to-be-optimized graphic text recognition result or a modified text of the graphic text recognition result, and a second text in each resource being a to-be-queried text corresponding to the first text in the resource; a query result obtained based on the to-be-queried text being better than a query result obtained based on the first text corresponding to the to-be-queried text;

[0008] if the target text is found in at least one resource, determining at least part of the second text contained in the at least one resource as a to-be-queried text corresponding to the text image.

[0009] In a possible implementation, the finding of the target text in the resource library comprises:

[0010] determining whether the initial text is Chinese;

[0011] if the initial text is Chinese, searching the initial text in the resource library;

[0012] if the initial text is not found in the resource library, modifying the initial text to obtain a modified text of the initial text;

[0013] searching the modified text of the initial text in the resource library.

[0014] In a possible implementation, the modifying the initial text comprises:

[0015] performing multiple modifications on the initial text to obtain multiple modified texts of the initial text;

[0016] each modified text is obtained by any of the following modifications: adding a wildcard at any position of the initial text, or replacing one word or multiple consecutive words in the initial text with a wildcard, or deleting one word or multiple consecutive words in the initial text.

[0017] In a possible implementation, the searching the target text in the resource library further comprises:

[0018] if the initial text is an English sentence, deleting a first word and a last word in the initial text to obtain a modified text of the initial text;

[0019] searching the modified text of the initial text in the resource library.

[0020] In a possible implementation, the process of determining the to-be-searched text comprises:

[0021] if the modified text of the initial text is a first text in multiple resources in the resource library, calculating a similarity between a first word and a last word in each of the multiple resources and the first word and the last word in the initial text;

[0022] determining a second text corresponding to a similarity that meets a condition as the to-be-searched text; the similarity that meets the condition is greater than the similarity that does not meet the condition.

[0023] In a possible implementation, the method further comprises:

[0024] if the initial text is an English word, searching the English word in a word library corresponding to a length of the English word;

[0025] if the English word is found, determining the English word as the to-be-searched text;

[0026] If not found, find an English word with the greatest similarity to the English word in a vocabulary corresponding to the target length as the text to be queried; the difference between the target length and the length of the English word is less than a threshold.

[0027] The second method of the present application provides a query method, comprising:

[0028] Obtaining a text image;

[0029] Obtaining a text to be queried corresponding to the text image according to the text acquisition method of the first aspect or any implementation manner of the first aspect;

[0030] Performing knowledge query based on the text to be queried to obtain a knowledge query result.

[0031] The third aspect of the present application provides a text acquisition device, comprising:

[0032] An image-text recognition module, configured to perform text recognition on the scanned text image to obtain an initial text;

[0033] A resource searching module, configured to search for a target text in a resource library, the target text being the initial text or a modified text obtained by modifying a partial text of the initial text; the resource library stores a plurality of resources, a first text in each resource being a to-be-optimized image-text recognition result or a modified text of the image-text recognition result, and a second text in each resource being a text to be queried corresponding to the first text in the resource; a query result obtained based on the text to be queried is better than a query result obtained based on the first text corresponding to the text to be queried; if the target text is found in at least one resource, at least part of the second text contained in the at least one resource is determined as a text to be queried corresponding to the text image.

[0034] The fourth aspect of the present application provides a query device, comprising:

[0035] An image acquisition module, configured to obtain a text image;

[0036] A query module, configured to obtain a text to be queried corresponding to the text image according to the text acquisition method of the first aspect or any implementation manner of the first aspect; and perform knowledge query based on the text to be queried to obtain a knowledge query result.

[0037] The fifth aspect of the present application provides a computer program product, comprising computer readable instructions, when the computer readable instructions run on an electronic device, causing the electronic device to implement the text acquisition method of the first aspect or any implementation manner of the first aspect, or implement the query method of the second aspect.

[0038] The sixth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected with the processor, wherein:

[0039] The memory is configured to store a computer program.

[0040] The processor is configured to execute the computer program, so that the electronic device can implement the text acquisition method of the first aspect or any implementation manner of the first aspect, or implement the query method of the second aspect.

[0041] The seventh aspect of the present application provides a computer storage medium, which carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the text acquisition method of the first aspect or any implementation manner of the first aspect, or implement the query method of the second aspect.

[0042] The eighth aspect of the present application provides a dictionary pen, comprising: a dictionary pen body, a processor arranged inside the dictionary pen body, an image acquisition module arranged on the surface of the dictionary pen body;

[0043] The image acquisition module is configured to acquire images.

[0044] The processor is configured to process the text image scanned by the image acquisition module, implement the text acquisition method of the first aspect or any implementation manner of the first aspect, or implement the query method of the second aspect.

[0045] By the above technical solution, the text acquisition method, the query method, the device, the equipment, the storage medium and the program product provided by the present application pre-construct a resource library, the first text in each resource in the resource library is the to-be-optimized image-text recognition result or the modified text of the image-text recognition result, and the second text in each resource is the to-be-queried text set according to the first text in the corresponding resource; the query result obtained by knowledge query based on the to-be-queried text is better than the query result obtained by knowledge query based on the first text corresponding to the to-be-queried text; based on the resource library, after text recognition is performed on the scanned text image to obtain a text recognition result, instead of taking the text recognition result as the to-be-queried text, the text recognition result is taken as an initial text, and a resource containing the initial text or a modified text thereof is searched in the resource library, and at least part of the second text in the found at least one resource is determined as the to-be-queried text corresponding to the text image; since the query result obtained by knowledge query based on the to-be-queried text is better than the query result obtained by knowledge query based on the first text corresponding to the to-be-queried text, therefore, the result of knowledge query based on the to-be-queried text corresponding to the text image is better than the query result obtained by knowledge query based on the initial text, and the information query quality in the image-text recognition scene is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0047] Figure 1 A flowchart of an implementation of the text acquisition method provided in this application;

[0048] Figure 2 A flowchart for searching for target text in a resource library provided by this application;

[0049] Figure 3 A flowchart for searching for target text in a resource library provided by this application;

[0050] Figure 4 A structural diagram of a text acquisition device provided by this application;

[0051] Figure 5 A schematic diagram of the structure of the dictionary pen provided in this application;

[0052] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION

[0053] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.

[0054] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0055] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0056] In order to improve the information query quality in the scenario of image-text recognition, the scheme is provided.

[0057] As shown in Figure 1 An implementation flowchart of the text acquisition method provided by the embodiment of the present application can include the following steps.

[0058] Step S101: performing text recognition on the scanned text image to obtain initial text.

[0059] When a user reads a paper file (including but not limited to a paper book) or an electronic file (including but not limited to an electronic book, a word document, etc.), if the user wants to query the information in the file (the paper file or the electronic file), the user can hold an electronic device (such as a dictionary pen) with a scanning function to scan the text in the file that needs to be queried to obtain a text image.

[0060] In order to perform text query, the dictionary pen needs to perform text recognition (such as optical character recognition: OCR) on the text image to obtain a text recognition result (which can also be referred to as an image-text recognition result), which is recorded as initial text for convenience of description and differentiation.

[0061] Due to the interference information in the scanned text in the file, or the error in text recognition, or the error caused by human operation, and various factors, various situations can occur in the text recognition result (i.e., the initial text). As shown in Table 1, some examples of the text recognition result provided by the embodiment of the present application are shown.

[0062] Table 1

[0063] The "text type" in Table 1 is some situations that can occur in the recognition result of text recognition on the text image, and the "use case" is an example corresponding to the "text type".

[0064] As can be seen from Table 1, some image-text recognition results directly as the text to be queried can affect the query effect, such as some image-text recognition results can have interference information (for example, numbers, characters, annotation marks, etc. in ancient poems or idioms), the image-text recognition result can also be incomplete information (for example, missing parts of words in poems or idioms, etc.), the image-text recognition result can also be misrecognized (for example, missing letters, adding letters, misrecognition, etc. in words), and directly using the image-text recognition result with the above situations as the text to be queried for information query can not query information, or the queried information is incorrect, etc., which reduces the user experience.

[0065] Of course, some OCR results directly as the query text will not affect the query effect, for example, the OCR result is the complete knowledge point name "countdown", "direct ratio", "high order equation" and so on.

[0066] Step S102: find the target text in the resource library, the target text is the initial text or the modified text obtained by modifying the partial text of the initial text.

[0067] The resource library stores a plurality of resources, the first text in each resource is an OCR result to be optimized or a modified text of the OCR result, and the second text in each resource is a query text set corresponding to the first text in the resource; the query result obtained based on the query text is better than the query result obtained by querying knowledge based on the first text corresponding to the query text.

[0068] In this application, each resource in the resource library is a text pair (including the first text and the second text), and the first text in the text pair is an OCR result to be optimized, that is, the first text is a possible text recognition result obtained by performing text recognition on a certain text image. It is possible that no knowledge can be queried or the queried knowledge is incorrect when the possible text recognition result is used as a query text to query knowledge; the second text in the text pair is a query text set in advance corresponding to the first text, and the quality of the query result obtained based on the query text is better than the quality of the query result obtained by querying knowledge based on the first text in the same text pair.

[0069] The first text and the second text in the same text pair satisfy a preset fuzzy matching rule. That is, the second text in each text pair is determined based on the first text in the text pair and the preset fuzzy matching rule.

[0070] As shown in Table 2, examples of fuzzy matching rules set for each type of OCR result use case provided by the embodiments of the present application are shown:

[0071] Table 2

[0072] Some fuzzy matching rules in Table 2 will be explained as follows:

[0073] For the text recognition result (i.e. the first text) of the type of "title + number", "sentence + number", "remove the number, hit the title or sentence" means removing the number in this type of text recognition result, and the obtained title or sentence is used as the corresponding second text.

[0074] For text recognition results of the "title + symbol" or "sentence + symbol" type (i.e., the first text), "remove symbols and hit the title or sentence" means removing the symbols (characters, annotation subscripts, etc.) in this type of text recognition results and using the resulting title or sentence as the corresponding second text.

[0075] For text recognition results (ie, first text) of the "title + other text" type, such as "title + sentence", "title + author", "title + dynasty + author", etc., the title in the first text is used as the corresponding second text.

[0076] For an incomplete verse whose length is greater than two Chinese characters as a result of text recognition, the complete verse corresponding to the incomplete verse is used as the corresponding second text.

[0077] If the text recognition result (i.e., the first text) includes two verses and one complete verse (for example, "A year is gone with the sound of firecrackers, and the spring breeze arrives"), the title of the ancient poem to which the verse belongs (i.e., "New Year's Day") is used as the corresponding second text.

[0078] For the case where the text recognition result (i.e., the first text) includes two verses but does not include a complete verse (for example: One year is gone with the sound of firecrackers, and the spring breeze brings warmth into Tu Su), the complete sentence corresponding to the text recognition result (i.e., One year is gone with the sound of firecrackers, and the spring breeze brings warmth into Tu Su) is used as the corresponding second text.

[0079] If the text recognition result (ie, the first text) is a complete scan of an ancient poem but misses one sentence, the title or the first sentence of the poem is used as the corresponding second text.

[0080] For text recognition results of the type "idiom + symbol" or "idiom missing + symbol" (i.e., the first text), the complete idiom is directly used as the corresponding second text.

[0081] The above examples use ancient poems and idioms as examples. Based on the above fuzzy matching rules, you can find the title of a poem based on the text recognition results of the "title + other text" type, find the title of a poem based on the verse, find the complete verse based on the fragment (i.e., incomplete verse), find the complete idiom based on the idiom with missing words, and so on.

[0082] For the 2-4 character mathematical knowledge point name, and the "4-character + number" mathematical knowledge point name, only the text recognition result (i.e. the first text) is the complete name of the knowledge point, which can match the complete name of the knowledge point, otherwise, it cannot match the complete name of the 2-4 character mathematical knowledge point, or it cannot match the complete name of the "4-character + number" mathematical knowledge point. That is, for the 2-4 character mathematical knowledge point name, and the complete name of the "4-character + number" mathematical knowledge point, the corresponding resource is the complete name of the mathematical knowledge point, or the first text and the second text of the text pair are the complete name of the same mathematical knowledge point.

[0083] For the case where the text recognition result (i.e. the first text) is "4 characters + symbol" type mathematical knowledge point, remove the symbol in the text recognition result to get the knowledge point name as the second text.

[0084] For mathematical knowledge points with more than 4 characters, the complete name of such mathematical knowledge points can be segmented, and at least 50% of the continuous multiple segments in the segmentation result are taken as the first text, and the complete name of the mathematical knowledge point is taken as the second text corresponding to the first text. When at least 50% of the continuous multiple segments in the segmentation result have multiple cases, multiple text pairs can be established, the first text in different text pairs is different continuous multiple segments in the segmentation result, and the continuous multiple segments account for more than or equal to 50% in the segmentation result; the second text in different text pairs is the complete name of the mathematical knowledge point. In this case, as long as the text recognition result hits at least 50% of the continuous content of the complete name of the digital knowledge point, the complete name of the digital knowledge point is taken as the query text.

[0085] Among them, for "n characters + symbol", "n characters + symbol + number" mathematical knowledge points, the complete name of such mathematical knowledge points can be segmented, and the symbols in the segmentation result are removed, the letters are not removed, to get the processed segmentation result, at least 50% of the continuous multiple segments in the processed segmentation result are taken as the first text, and the complete name of the mathematical knowledge point is taken as the second text corresponding to the first text. When at least 50% of the continuous multiple segments in the processed segmentation result have multiple cases, multiple text pairs can be established, the first text in different text pairs is different continuous multiple segments in the processed segmentation result, and the continuous multiple segments account for more than or equal to 50% in the segmentation result; the second text in different text pairs is the complete name of the mathematical knowledge point. In this case, as long as the text recognition result hits at least 50% of the continuous content in the processed segmentation result, the complete name of the digital knowledge point is taken as the query text.

[0086] For an English sentence, the graph recognition result of the sentence often has errors in the first word or the last word of the sentence. Based on this, a complete sentence in a textbook or a picture book with a length greater than a threshold value can be taken as a second text, and the first word and the last word of the complete sentence are removed to obtain a corresponding first text. A text pair is formed as a resource. Based on this, after obtaining a text recognition result, the first and last words of the text recognition result are removed, and the text recognition result after removing the first and last words is searched in a resource library. If the first text is found, the similarity between each second text corresponding to the first text and the text recognition result is calculated, and the second text with a similarity greater than a similarity threshold value is determined as the to-be-queried text.

[0087] Step S103: If the target text is found in at least one resource, at least part of the second text contained in the at least one resource is determined as the to-be-queried text corresponding to the text image.

[0088] That is, if the target text is the first text or the second text in at least one resource in the resource library, the second text in at least part of the at least one resource is determined as the to-be-queried text corresponding to the text image.

[0089] In the resource library, it is possible that only one resource contains the target text, or multiple resources contain the target text, or no resource contains the target text.

[0090] In the case where at least one (denoted as m, m is an integer greater than 0) resource contains the target text, the second text in the m resources can be taken as the to-be-queried text respectively, that is, m to-be-queried texts are obtained. Based on this, in the subsequent knowledge query, the knowledge query can be performed based on each to-be-queried text respectively, and m knowledge query results are obtained.

[0091] Alternatively, in the case where m is greater than 1, the k texts most similar to the initial text can be screened out from the second texts in the m resources as k to-be-queried texts. k is less than m. Based on this, in the subsequent knowledge query, the knowledge query can be performed based on each to-be-queried text respectively, and k knowledge query results are obtained.

[0092] The text acquisition method provided in the application pre-constructs a resource library, the first text in each resource in the resource library is a text recognition result to be optimized or a modified text of the text recognition result, and the second text in each resource is a to-be-queried text set according to the first text in the corresponding resource; the query result obtained based on the to-be-queried text is better than the query result obtained based on the first text corresponding to the to-be-queried text. Based on this, after a scanned text image is subjected to text recognition to obtain a text recognition result, instead of taking the text recognition result as a to-be-queried text, the text recognition result is taken as an initial text, a resource containing the initial text or a modified text thereof is searched for in the resource library, and the second text in at least part of the at least one resource found is determined as a to-be-queried text corresponding to the text image; since the query result obtained based on the to-be-queried text is better than the query result obtained based on the first text corresponding to the to-be-queried text, the query result obtained based on the to-be-queried text corresponding to the text image is better than the query result obtained based on the initial text, and the information query quality in the text recognition scenario is improved.

[0093] In an optional embodiment, an implementation flowchart for searching for a target text in a resource library as shown in Figure 2 may include:

[0094] Step S201: determining whether the initial text is Chinese.

[0095] Optionally, whether the initial text is Chinese can be determined based on the unicode code of at least part of the characters or words in the initial text.

[0096] As an example, if the unicode code of each character or word in the initial text is within a first preset range, it is determined that the initial text is Chinese; if the unicode code of each character or word in the initial text is within a second preset range, it is determined that the initial text is English. The first preset range and the second preset range are different.

[0097] As shown in Table 2, in the application, Chinese can include ancient poems, idioms, names of knowledge points in various disciplines, etc. English can include English words or English sentences in teaching materials, English words or English sentences in picture books, etc.

[0098] Step S202: if the initial text is Chinese, searching for the initial text in the resource library.

[0099] In the case that the initial text is in Chinese, the initial text can be directly searched in the resource library. In order to improve the searching speed, the hash value of the first text and the hash value of the second text can be recorded for each resource in the resource library. Accordingly, when searching for the initial text in the resource library, the hash value of the initial text can be calculated, and then the hash value of the initial text is searched in the resource library. If the hash value of the initial text is found to be the hash value of the first text in a certain resource, it is determined that the first text in the certain resource is the initial text. If the hash value of the initial text is found to be the hash value of the second text in a certain resource, it is determined that the second text in the certain resource is the initial text.

[0100] Optionally, if multiple resources containing the initial text are found, each second text in the found multiple resources can be taken as a to-be-queried text.

[0101] Optionally, if multiple resources containing the initial text are found, the semantic similarity of each second text in the multiple resources to the initial text can be calculated, and the second text with a semantic similarity to the initial text greater than a threshold value is determined as a to-be-queried text.

[0102] Step S203: If the initial text is not found in the resource library, the initial text is modified to obtain a modified text of the initial text.

[0103] If the initial text is not found in the resource library, i.e., the initial text is not included in any resource, the initial text can be modified locally. The modification of the initial text can include but is not limited to adding a word, deleting a word, replacing a word, etc.

[0104] Only one word in the initial text can be modified, or multiple consecutive words in the initial text can be modified. The number of modified words can be determined according to the length of the initial text and a preset modification ratio. As an example, the number of modified words can be obtained by multiplying the length of the initial text by the modification ratio and rounding up the product. For example, if the modification ratio is 10% and the length of the sentence is 10 (i.e., there are 10 words), only one word is modified. If the length of the sentence is less than 10, only one word is modified. If the length of the sentence is 22, three consecutive words are modified.

[0105] For any modification method, whether it is to modify one word or to modify multiple consecutive words, there are multiple cases. For example, replacing one word or deleting one word can replace or delete any one word in the initial text. If the length of the sentence is L, there are L replacement methods or L deletion methods, and L modified texts can be obtained. For another example, adding one word can be added between any two words in the initial text, or added before the beginning of the initial text, or added after the end of the initial text. If the length of the sentence is L, there can be L+1 modified texts.

[0106] Step S204: searching for the modified text of the initial text in the resource library.

[0107] Optionally, the hash value of the modified text can be calculated, and the hash value is searched in the resource library. If the hash value of the modified text is found to be the hash value of the first text in a certain resource, it is determined that the first text in the certain resource is the modified text of the initial text, and if the hash value of the modified text is found to be the hash value of the second text in a certain resource, it is determined that the second text in the certain resource is the modified text of the initial text.

[0108] In an optional embodiment, one implementation of the above-mentioned modification of the initial text can be as follows:

[0109] The initial text is modified in multiple ways to obtain multiple modified texts of the initial text.

[0110] Each modified text is obtained by modifying the initial text in any of the following ways: adding a wildcard at any position of the initial text, or replacing one word or multiple consecutive words in the initial text with a wildcard, or deleting one word or multiple consecutive words in the initial text, etc.

[0111] The wildcard can be a specified symbol, such as "\t", a question mark "?", or an asterisk "*" and the like.

[0112] In the case of adding or replacing words, by adding a wildcard or replacing it with a wildcard, it is not necessary to traverse all Chinese characters to modify the text, reducing the number of modified texts. Correspondingly, the first text in the resource library can also be a text with a wildcard, reducing the amount of data in the resource library, and further reducing the amount of data processing in the text acquisition process, improving the text acquisition efficiency.

[0113] In an optional embodiment, in the case where the initial text is an English sentence, one implementation flowchart of the above-mentioned searching for the target text in the resource library can be as shown in FIG. 1C, which can include: Figure 3

[0114] Step S301: deleting the first word and the last word of the initial text to obtain the modified text of the initial text.

[0115] Optionally, before performing step S301, the length of the English sentence can be determined first, and if the length of the sentence is greater than a length threshold, step S301 is performed.

[0116] If the length of the English sentence is less than or equal to the length threshold, it can be considered that the sentence is correct, and the initial text can be directly used as the text to be queried.

[0117] ​​Step S302: finding the modified text of the initial text in the resource library.

[0118] Further, in the case that the initial text is an English sentence, the process of determining the text to be queried can include:

[0119] If the modified text of the initial text is the first text in one of the resources in the resource library, the second text in the resource is determined as the text to be queried.

[0120] If the modified text of the initial text is the first text in a plurality of resources (denoted as m, m is an integer greater than 1) in the resource library, the similarity of the first and last words in the second text in each of the m resources and the first and last words in the initial text can be calculated.

[0121] Optionally, to calculate the similarity of word B and word A, the proportion of the same letters in the same positions in word B can be used to represent the similarity of the two words. The greater the proportion of the same letters in the same positions in the target word, the greater the similarity of the two words. The target word is the non-reference word in the two words. For example, to calculate the similarity of word B and word A, word A is determined as the reference word and word B is determined as the non-reference word.

[0122] In the embodiments of the present application, in the case of calculating the similarity of the first word of the second text in the resource and the first word of the initial text, the first word of the second text in the resource is taken as the target word; in the case of calculating the similarity of the last word of the second text in the resource and the last word of the initial text, the last word of the second text in the resource is taken as the target word.

[0123] As an example, since the first word of a sentence may miss the first letter of the word, based on this, in the case of calculating the similarity of the first word of the second text in the resource and the first word of the initial text, the comparison can be started from the last letter of the two first words, to determine whether the letters in the same positions of the two first words are the same, to count the number of the same letters in the same positions of the two first words, and to determine the proportion of the counted number in the first word of the second text in the resource as the similarity of the first word of the second text in the resource and the first word of the initial text.

[0124] For example, assuming that the first word in the second text is "Grapes" and the first word in the initial text is "Orapes", the number of identical letters in the same position in the two words is counted from the tail letter of the two first words. In this example, from the tail letter "s" of the two first words, the identical letters in the same position are determined in sequence as: s, e, p, a, and r, which are 5 letters in total. The length of the first word "Grapes" in the second text is 6, and thus the ratio is 5 / 6=83%, i.e., the similarity is 83%.

[0125] For example, assuming that the first word in the second text is "Grapes" and the first word in the initial text is "Orapes", the number of identical letters in the same position in the two words is counted from the tail letter of the two first words. In this example, from the tail letter "s" of the two first words, the identical letters in the same position are determined in sequence as: s, e, p, a, and r, which are 5 letters in total. The length of the first word "Grapes" in the second text is 6, and thus the ratio is 5 / 6=83%, i.e., the similarity is 83%.

[0126] For example, assuming that the first word in the second text is "Grapes" and the first word in the initial text is "Orapes", the number of identical letters in the same position in the two words is counted from the tail letter of the two first words. In this example, from the tail letter "s" of the two first words, the identical letters in the same position are determined in sequence as: s, e, p, a, and r, which are 5 letters in total. The length of the first word "Grapes" in the second text is 6, and thus the ratio is 5 / 6=83%, i.e., the similarity is 83%.

[0127] For example, assuming that the first word in the second text is "Grapes" and the first word in the initial text is "Orapes", the number of identical letters in the same position in the two words is counted from the tail letter of the two first words. In this example, from the tail letter "s" of the two first words, the identical letters in the same position are determined in sequence as: s, e, p, a, and r, which are 5 letters in total. The length of the first word "Grapes" in the second text is 6, and thus the ratio is 5 / 6=83%, i.e., the similarity is 83%.

[0128] For example, if the last word of the second text is "small" and the first word of the initial text is "smaall", the number of same letters at the same position in the two words is counted from the first letter of the two last words, in this example, from the first letter "s" of the two last words, the same letters at the same position are determined as follows: s, m, a, and 1, which are four letters. The length of the last word "small" of the second text is five, and thus the ratio is 4 / 5=80%, i.e., the similarity is 80%.

[0129] After the similarity of the first word and the similarity of the last word are calculated, the second text whose similarity of the first and last words meets the condition is selected as the query text according to a preset screening strategy.

[0130] Optionally, the second text whose similarity of the first word of the initial text is greater than the similarity threshold and the similarity of the last word of the initial text is greater than the similarity threshold is selected as the query text.

[0131] Optionally, if there is no second text whose similarity of the first and last words of the initial text is greater than the similarity threshold, the second text whose similarity of the first word of the initial text is greater than the similarity threshold and the second text whose similarity of the last word of the initial text is greater than the similarity threshold are selected as the query text.

[0132] In an optional embodiment, if the initial text is an English word, the English word can be searched in a word library corresponding to the length of the English word. Different lengths correspond to different word libraries, and the lengths of the words in the same word library are the same.

[0133] That is, the present application divides the words into different word libraries according to the length of the words, and in the case of an initial text being an English word, the English word is searched in the words with the same length as the English word.

[0134] If the English word is found in the word library corresponding to the length of the English word, it means that the English word is a complete word, and the English word can be used as the query text.

[0135] If the English word is not found in the word library corresponding to the length of the English word, the English word with the greatest similarity to the English word can be searched in the word library with a length close to the length of the English word as the query text.

[0136] Optionally, the English word with the greatest similarity to the English word can be searched in the word library corresponding to the target length. The target length is less than the threshold value from the length of the English word.

[0137] For example, if the length of the English word as the initial text is s, the target length can include s+1, s and s-1. Correspondingly, the word library corresponding to the target length has three, i.e., the word library corresponding to the length s+1, the word library corresponding to the length s, and the word library corresponding to the length s-1.

[0138] Alternatively, the target length can be s+2, s+1, s, s-1 and s-2. Correspondingly, the word library corresponding to the target length has five, i.e., the word library corresponding to the length s+2, the word library corresponding to the length s+1, the word library corresponding to the length s, the word library corresponding to the length s-1, and the word library corresponding to the length s-2.

[0139] The present application finds that there are many hash conflicts for English words, and therefore, it is not suitable for hash matching. The present application uses the edit distance to represent the similarity between English words. The smaller the edit distance between two English words, the greater the similarity between the two English words.

[0140] As an example, the edit distance can use the Levenshtein distance, which is an algorithm for measuring the difference between two strings. It calculates the minimum number of editing operations required to convert one string (e.g., an English word) to another string (e.g., an English word). These editing operations include inserting, deleting and replacing characters.

[0141] The core idea of the Levenshtein distance is to calculate the minimum edit distance between two strings by dynamic programming method. Specifically, it constructs a two-dimensional matrix, where each cell represents the minimum number of editing operations required to convert the first i characters of one string to the first j characters of another string. For specific implementation, refer to the existing implementation, which will not be described here.

[0142] In the case where there are multiple English words in the word library that are most similar to the scanned English word (i.e., the initial text), the word that is most similar to the scanned English word among the words that are most similar to the scanned English word can be used as the text to be queried.

[0143] As shown in Table 3, an example of the matching result when the text recognition result is an English word is provided by the embodiment of the present application.

[0144] Table 3

[0145]

[0146] Corresponding to the foregoing text acquisition method embodiment, the present application also provides a query method. The query method provided by the embodiment of the present application can include:

[0147] Obtain a text image, which can be obtained by the electronic device through real-time scanning by an image acquisition module, or can be read from a storage unit.

[0148] Process the text image according to the text acquisition method as described above to obtain a to-be-queried text corresponding to the text image. The to-be-queried text corresponding to the text image can be one or more.

[0149] Perform knowledge query based on the to-be-queried text corresponding to the text image to obtain a knowledge query result. That is, the to-be-queried text is used as a search keyword to perform knowledge query. In the case where the to-be-queried text corresponding to the text image is more than one, knowledge query is performed based on each to-be-queried text to obtain a plurality of knowledge query results, each to-be-queried result corresponding to the first to-be-queried text.

[0150] The present application improves the information query quality in the scenario of image-text recognition.

[0151] Corresponding to the method embodiment, the present application also provides a text acquisition device. A structural diagram of the text acquisition device provided by the present application embodiment is shown in Figure 4 may include:

[0152] An image-text recognition module 401 and a resource finding module 402;

[0153] The image-text recognition module 401 is configured to perform text recognition on the scanned text image to obtain an initial text.

[0154] The resource finding module 402 is configured to find a target text in a resource library, the target text being the initial text or a modified text obtained by modifying a partial text of the initial text. The resource library stores a plurality of resources, a first text in each resource being a to-be-optimized image-text recognition result or a modified text of the image-text recognition result, and a second text in each resource being a to-be-queried text corresponding to the first text in the resource. A query result obtained based on the to-be-queried text is better than a query result obtained based on the first text corresponding to the to-be-queried text. If the target text is found in at least one resource, at least part of the second text contained in the at least one resource is determined as a to-be-queried text corresponding to the text image.

[0155] The modules in the text acquisition device described above can be all or partially implemented by software, hardware, and a combination thereof. The modules described above can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to the modules.

[0156] The text acquisition device provided in the application is preconfigured with a resource library, the first text in each resource in the resource library is a to-be-optimized OCR result or a modified text of the OCR result, and the second text in each resource is a to-be-queried text set according to the first text in the resource; the query result obtained based on the to-be-queried text is better than the query result obtained based on the first text corresponding to the to-be-queried text, based on the resource library, after a scanned text image is subjected to text recognition to obtain a text recognition result, instead of taking the text recognition result as the to-be-queried text, the text recognition result is taken as an initial text, a resource containing the initial text or a modified text thereof is searched for in the resource library, and the second text in at least part of the at least one resource found is determined as the to-be-queried text corresponding to the text image; since the query result obtained based on the to-be-queried text is better than the query result obtained based on the first text corresponding to the to-be-queried text, therefore, the result of knowledge query based on the to-be-queried text corresponding to the text image is better than the query result obtained based on the initial text, and the information query quality in the OCR scenario is improved.

[0157] The detailed functions and extended functions of the OCR module 401 and the resource searching module 402 can refer to the foregoing method embodiments, and will not be described here again.

[0158] Corresponding to the method embodiments, the application further provides a query device, and the query device provided in the application can include:

[0159] An image acquisition module, configured to obtain a text image;

[0160] A query module, configured to obtain a to-be-queried text corresponding to the text image according to the foregoing text acquisition method; and perform knowledge query based on the to-be-queried text to obtain a knowledge query result.

[0161] As shown in Figure 5 , a structural schematic diagram of a dictionary pen provided in an embodiment of the application can include:

[0162] A dictionary pen body 501, an image acquisition module 502 arranged on the surface of the dictionary pen body, and a processor 503 arranged inside the dictionary pen body 501;

[0163] The dictionary pen body 501 is configured to be held by a user.

[0164] The image acquisition module 502 is configured to acquire images. Specifically, the image acquisition module 502 can acquire image information of a text region in real time when a user holds the dictionary pen to scan the text region, including the text and the background content around the text. The image acquisition module can support multiple adjustment modes, such as focal length adjustment and optical zoom, to adapt to the text acquisition requirements in different font sizes and complex backgrounds.

[0165] The processor 503 is configured to process the text image scanned by the image acquisition module according to any of the text acquisition methods or query methods provided in the embodiments of the present application.

[0166] The dictionary pen provided in the present application is pre-configured with a resource library. The first text in each resource in the resource library is a text recognition result to be optimized or a modified text of the text recognition result, and the second text in each resource is a text to be queried corresponding to the first text in the resource. The query result obtained based on the text to be queried is better than the query result obtained based on the first text corresponding to the text to be queried. Based on the resource library, after the text recognition result of the scanned text image is obtained by text recognition, the dictionary pen does not take the text recognition result as the text to be queried, but takes the text recognition result as the initial text, and searches for a resource containing the initial text or the modified text of the initial text in the resource library. The second text in at least part of the at least one resource found is determined as the text to be queried corresponding to the text image. Since the query result obtained based on the text to be queried is better than the query result obtained based on the first text corresponding to the text to be queried, the query result obtained based on the text to be queried corresponding to the text image is better than the query result obtained based on the initial text, and the information query quality of the dictionary pen is improved.

[0167] Optionally, the refinement function and the extension function of the processor 503 can be implemented by referring to the foregoing method embodiments, which will not be described herein.

[0168] Further, the dictionary pen can further include a display module arranged on the surface of the dictionary pen body 501, configured to display the knowledge queried based on the text to be queried. As an example, the display module can include, but is not limited to, any of the following display modules: an LCD (Liquid Crystal Display) display module, an OLED (Organic Light Emitting Display) display module, an electronic ink screen (E-ink), and the like.

[0169] In the embodiments of the present application, an electronic device is further provided. Referring to Figure 6, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application can be a terminal device (for example, this dictionary pen, or other electronic devices that establish a communication connection with the dictionary pen, etc.), or a server (which can be a single server, a server cluster, or a cloud server, etc.). Figure 6 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0170] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0171] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0172] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any text acquisition method or query method provided in the embodiment of the present application.

[0173] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any text acquisition method or query method provided in the embodiment of the present application.

[0174] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate units can or can not be physically separate, and the units shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be realized by means of software and the necessary general hardware, and of course can also be realized by special hardware including special integrated circuits, special CPUs, special memories, special components and the like. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and the specific hardware structure for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit. However, for the present application, the software program implementation is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for making a computer device (which can be a personal computer, training device, or network device, etc.) execute the methods described in each embodiment of the present application.

[0176] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, it can be realized in the form of computer program product in whole or in part. Professional technicians can use different methods to realize the described functions for each specific solution, but such implementation should not be considered beyond the scope of the present application.

[0177] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or data storage device such as training device, data center, etc. integrated with one or more available media sets. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, Solid State Disk (SSD)) and the like.

[0178] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0179] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text acquisition method, characterized in that: include: Perform text recognition on the scanned text image to obtain the initial text; Searching for a target text in a resource library, where the target text is the initial text or a modified text obtained by modifying a portion of the initial text; the resource library stores a plurality of resources, wherein the first text in each resource is a graphic-text recognition result to be optimized or a modified text of the graphic-text recognition result, and the second text in each resource is a text to be queried set corresponding to the first text in the resource; A query result obtained by performing a knowledge query based on the text to be queried is better than a query result obtained by performing a knowledge query based on the first text corresponding to the text to be queried; If the target text is found in at least one resource, at least a portion of the second text contained in the at least one resource is determined as the text to be queried corresponding to the text image.

2. The method according to claim 1, characterized in that The step of searching for the target text in the resource library includes: Determining whether the initial text is in Chinese; If the initial text is in Chinese, searching for the initial text in the resource library; If the initial text is not found in the resource library, modify the initial text to obtain a modified text of the initial text; Searching the resource library for a modified text of the initial text.

3. The method according to claim 2, characterized in that The modifying of the initial text includes: Performing multiple modifications on the initial text to obtain multiple modified texts of the initial text; Each modified text is modified by any of the following methods: adding a wildcard at any position of the original text, or replacing one or more consecutive words in the original text with a wildcard, or deleting one or more consecutive words in the original text.

4. The method according to claim 2, characterized in that The searching for the target text in the resource library further includes: If the initial text is an English sentence, deleting the first word and the last word in the initial text to obtain a modified text of the initial text; Searching the resource library for a modified text of the initial text.

5. The method according to claim 4, characterized in that The process of determining the text to be queried includes: If the modified text of the initial text is the first text in the plurality of resources in the resource library, calculating the similarity between the first and last words of the second text in each of the plurality of resources and the first and last words in the initial text; The second text corresponding to the similarity that meets the condition is determined as the text to be queried; the similarity that meets the condition is greater than the similarity that does not meet the condition.

6. The method according to claim 2, characterized in that Also includes: If the initial text is an English word, searching for the English word in a vocabulary corresponding to the length of the English word; If found, the English word is determined as the text to be queried; If the search result is not found, searching for an English word with the greatest similarity to the English word in the vocabulary corresponding to the target length as the search result; The difference between the target length and the length of the English word is smaller than a threshold.

7. A query method, characterized in that: include: Get text image; Processing the text image according to the text acquisition method according to any one of claims 1 to 6 to obtain the text to be queried corresponding to the text image; A knowledge query is performed based on the text to be queried to obtain a knowledge query result.

8. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the text acquisition method according to any one of claims 1 to 6, or implement the query method according to claim 7.

9. An electronic device, characterized in that: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program so that the electronic device can implement the text acquisition method according to any one of claims 1 to 6, or implement the query method according to claim 7.

10. A computer storage medium, characterized in that The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the text acquisition method as described in any one of claims 1 to 6, or the query method as described in claim 7.

11. A dictionary pen, characterized in that: The dictionary pen comprises: a dictionary pen body, a processor arranged inside the dictionary pen body, and an image acquisition module arranged on the surface of the dictionary pen body; The image acquisition module is used to acquire images; The processor is configured to: The text image scanned by the image acquisition module is processed to implement the text acquisition method according to any one of claims 1 to 6, or to implement the query method according to claim 7.