A method, device and storage medium for intelligent text recognition of scanned images

Through image recognition processing and non-related content deletion of point reading devices, the problem of inconsistent scanning content of point reading devices is solved, and an efficient learning experience for users to obtain target content without frequent operations is achieved.

CN114596568BActive Publication Date: 2025-08-22SUZHOU QINGRUI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111652173.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-08-22
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The content scanned by the click-reading device is inconsistent with the content that the user actually wants, resulting in frequent operations to obtain the required content, and the user experience is poor.

Method used

By performing image recognition processing on the target image obtained by the point reading device, the image content is determined, and non-related content is deleted to obtain the target content that the user wants to learn.

Benefits of technology

This reduces the frequency of manual selection by users and improves the user experience of point reading devices, so that users can directly obtain the required content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596568B_ABST
    Figure CN114596568B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, and storage medium for intelligent text recognition of scanned images, relating to the technical field of educational electronic equipment. The method addresses the issues of frequent operations and poor user experience when users use point-reading devices for learning. The specific solution includes: acquiring a target image in response to a user's scanning operation using the point-reading device; performing image recognition processing on the target image to determine the image content of the target image, where the image content includes the content obtained after the user scans the image using the point-reading device; and deleting irrelevant content from the image content to obtain the target content, where the target content is the content to be learned by the user, and irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of educational electronic equipment, and in particular to an intelligent text recognition method, device and storage medium for scanned images. Background Art

[0002] Point-reading devices can acquire images by scanning, and then perform image-to-text conversion, speech synthesis, and other processing on the scanned images, and output voice data for users to learn.

[0003] However, in the related art, the content scanned by the point reading device may not be consistent with the content the user actually wants to read, making the voice data output by the point reading device unable to meet the user's needs. The user needs to frequently use the point reading device to scan to obtain the desired content, resulting in a poor user experience. Summary of the Invention

[0004] The present invention provides a method, device and storage medium for intelligent text recognition of scanned images, which solves the problem of frequent operations and poor user experience when users use point reading devices for learning.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a method for intelligent text recognition of a scanned image, the method comprising:

[0007] In response to a scanning operation of a user using a point reading device, acquiring a target image;

[0008] Performing image recognition processing on the target image to determine the image content of the target image, wherein the image content includes the content obtained after the user scans the target image using a point reading device;

[0009] Delete irrelevant content in the image content to obtain target content, where the target content is the content to be learned by the user in the image content, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering conditions.

[0010] The intelligent text recognition method for scanned images provided by the present invention performs image recognition processing on a target image captured by a point-reading device, determines the image content of the target image, and deletes irrelevant content from the image to obtain the target content to be learned by the user. This method allows direct access to the target content without the user having to manually select it again. Compared to related art techniques that require users to frequently scan with a point-reading device to obtain the desired content, the intelligent text recognition method for scanned images provided by the embodiments of the present invention can solve the problem of frequent operations and poor user experience when users use point-reading devices for learning.

[0011] In a possible implementation, the image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture.

[0012] In a possible implementation, when the image content includes text content, deleting irrelevant content in the image content to obtain the target content includes:

[0013] Preprocessing the image content to obtain the words included in the processed content;

[0014] Irrelevant words in the words included in the processed content are deleted to obtain the target content.

[0015] In a possible implementation, the preprocessing of the image content to obtain the words included in the processed content includes:

[0016] Delete redundant information of image content to obtain valid content;

[0017] determining at least one segmentation point of the valid content, where the segmentation point is used to segment adjacent words in the valid content;

[0018] The valid content is segmented based on at least one segmentation point to obtain words included in the processed content.

[0019] In a possible implementation, the non-relevant words in the words included in the content after the above-mentioned deletion processing are obtained to obtain the target content, including:

[0020] When the processed content includes multiple words, deleting incomplete words from the multiple words to obtain at least one complete word;

[0021] The target content is determined based on at least one complete word.

[0022] In a possible implementation, determining the target content based on at least one complete word includes:

[0023] If at least one complete word satisfies a preset rule, then it is determined that the target content includes at least one complete word, and the preset rule includes at least one of the following: a sentence, a phrase, or a proper noun;

[0024] If at least one complete word does not meet the preset rules, a preset evaluation strategy is adopted to determine the evaluation parameters of each complete word in the at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

[0025] In one possible implementation, determining the target content based on at least one evaluation parameter includes:

[0026] Determining a target evaluation parameter that satisfies the parameter rule among at least one evaluation parameter;

[0027] If the target evaluation parameter is a parameter, the complete word corresponding to the target evaluation parameter is determined as the target content;

[0028] If the target evaluation parameter includes multiple parameters, multiple complete words corresponding to the target evaluation parameter are obtained, and the word scanned first by the point reading device among the multiple complete words is determined as the target content.

[0029] In a second aspect, the present invention provides an intelligent text recognition device for a scanned image, the intelligent text recognition device for a scanned image comprising:

[0030] an acquisition unit, configured to acquire a target image in response to a scanning operation of a user using a point reading device;

[0031] The processing unit is used to perform image recognition processing on the target image acquired by the acquisition unit, determine the image content included in the target image, which includes the content obtained after the user uses the point reading device to scan, and delete irrelevant content in the image content to obtain the target content, which is the content to be learned by the user in the image content, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering conditions.

[0032] In a possible implementation, the image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture.

[0033] In a possible implementation, when the image content includes content in text format, the processing unit is specifically configured to:

[0034] Preprocessing the image content to obtain the words included in the processed content;

[0035] Irrelevant words in the words included in the processed content are deleted to obtain the target content.

[0036] In a possible implementation, the processing unit is specifically configured to:

[0037] Delete redundant information of image content to obtain valid content;

[0038] Determine at least one segmentation point of the valid content, where the segmentation point is used to segment adjacent words in the valid content;

[0039] The valid content is segmented based on at least one segmentation point to obtain words included in the processed content.

[0040] In a possible implementation, the processing unit is specifically configured to:

[0041] When the processed content includes multiple words, deleting incomplete words from the multiple words to obtain at least one complete word;

[0042] The target content is determined based on at least one complete word.

[0043] In a possible implementation, the processing unit is specifically configured to:

[0044] If at least one complete word satisfies a preset rule, then it is determined that the target content includes at least one complete word, and the preset rule includes at least one of the following: a sentence, a phrase, or a proper noun;

[0045] If at least one complete word does not meet the preset rules, a preset evaluation strategy is adopted to determine the evaluation parameters of each complete word in the at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

[0046] In a possible implementation, the processing unit is specifically configured to:

[0047] Determining a target evaluation parameter that satisfies the parameter rule among at least one evaluation parameter;

[0048] If the target evaluation parameter is a parameter, the complete word corresponding to the target evaluation parameter is determined as the target content;

[0049] If the target evaluation parameter includes multiple parameters, multiple complete words corresponding to the target evaluation parameter are obtained, and the word scanned first by the point reading device among the multiple complete words is determined as the target content.

[0050] In a third aspect, the present invention provides an intelligent text recognition device for scanned images, comprising a processor and a memory. The memory is configured to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the intelligent text recognition device for scanned images performs the intelligent text recognition method for scanned images according to the first aspect and any possible implementation thereof.

[0051] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed on an intelligent text recognition device for scanned images, the intelligent text recognition device for scanned images executes an intelligent text recognition method for scanned images as described in the first aspect or any one of the possible implementations of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of the structure of an intelligent text recognition system provided by an embodiment of the present invention;

[0053] Figure 2 This is a schematic diagram of the structure of an intelligent text recognition device for scanned images provided by an embodiment of the present invention;

[0054] Figure 3 One of the flow charts of the intelligent text recognition method for scanned images provided by an embodiment of the present invention;

[0055] Figure 4 The second flowchart of the intelligent text recognition method for scanned images provided by the embodiment of the present invention;

[0056] Figure 5 Flowchart 3 of the intelligent text recognition method for scanned images provided by an embodiment of the present invention;

[0057] Figure 6 One of the interface diagrams of the intelligent recognition application provided by an embodiment of the present invention;

[0058] Figure 7 This is a second schematic diagram of the interface of the intelligent recognition application provided by an embodiment of the present invention;

[0059] Figure 8 This is a second structural diagram of the intelligent text recognition device for scanned images provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0061] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.

[0062] Additionally, the use of “based on” or “according to” is meant to be open and inclusive, as a process, step, calculation, or other action “based on” or “according to” one or more stated conditions or values ​​may, in practice, be based on additional conditions or values ​​beyond those stated.

[0063] To address the frequent operations and poor user experience associated with using point-reading devices for learning, embodiments of the present invention provide a method, device, and storage medium for intelligent text recognition of scanned images. This method performs image recognition processing on a target image captured by the point-reading device, determines the image content of the target image, and removes irrelevant content from the image to obtain the target content to be learned. This allows users to directly access the target content without having to manually select it again. After processing the target content, relevant information corresponding to the target content is obtained, facilitating learning.

[0064] The intelligent text recognition method for scanned images provided in the embodiments of the present invention is implemented by an intelligent text recognition device for scanned images. The intelligent text recognition device for scanned images can be a point-reading device or a processor of the point-reading device. The embodiments of the present invention use the point-reading device executing the intelligent text recognition method for scanned images as an example. For example, the point-reading device can be a point-reading pen or a scanning pen.

[0065] In one scenario, when a point-reading device executes the intelligent text recognition method for scanned images provided by an embodiment of the present invention, after acquiring a target image, it can perform image recognition processing to determine the image content included in the target image. The point-reading device can also delete irrelevant content from the image content to obtain the target content. The point-reading device can also process the target content to obtain relevant information about the target content. The image content includes the content obtained after the user uses the point-reading device to scan, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior trigger conditions, that is, the irrelevant content is the content that the user does not want to pay attention to. The relevant information may include any of the following: part of speech, word meaning, phonetic symbols, example sentences, voice data, and translation.

[0066] In another scenario, the image recognition processing, deletion of irrelevant content in the image content, and processing of the target content can be provided by the server. Specifically, the intelligent text recognition method for scanned images provided by the embodiment of the present invention can be applied to an intelligent text recognition system. Figure 1 FIG. 1 shows a schematic diagram of the structure of an intelligent text recognition system provided by an embodiment of the present invention. Figure 1 As shown, the intelligent text recognition system may include: a point reading device 11 and a server. The point reading device 11 is connected to the server via wired communication or wireless communication.

[0067] In some embodiments, the server may be a single server, a server cluster, or a cloud computing service platform. Figure 1 In the figure, a server cluster including one server is taken as an example.

[0068] The point reading device 11 is used to respond to the user's scanning operation using the point reading device 11, obtain the target image, send a processing request to the server 12, the processing request including the target image, and receive the target content and related information corresponding to the target content sent by the server 12.

[0069] The server 12 is used to receive the processing request sent by the point reading device 11, perform image recognition processing on the target image in the processing request, obtain the image content, delete irrelevant content in the image content, obtain the target content, and obtain relevant information of the target content, and return the target content and relevant information corresponding to the target content to the point reading device 11.

[0070] The basic hardware structure of the point reading device 11 and the server is similar, both of which include Figure 2 The components of the intelligent text recognition device for scanning images are shown below. Figure 2 Taking the intelligent text recognition device for scanned images as an example, the hardware structure of the point reading device and the server is introduced.

[0071] like Figure 2 As shown, the intelligent text recognition device for scanned images may include: a processor 21, a memory 22, a communication interface 23 and a bus 24. The processor 21, the memory 22 and the communication interface 23 may be connected via the communication bus 24.

[0072] The processor 21 is the control center of the intelligent text recognition device for scanning images. It can be a single processor 21 or a collective term for multiple processing elements. For example, the processor 21 can be a general-purpose central processing unit (CPU) or other general-purpose processors 21. The general-purpose processor 21 can be a microprocessor 21 or any other conventional processor 21.

[0073] As an embodiment, the processor 21 may include one or more CPUs, for example, Figure 2 CPU0 and CPU1 are shown.

[0074] The memory 22 may be a read-only memory 22 (ROM) or other type of static storage device that can store static information and instructions, a random access memory 22 (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory 22 (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0075] In one possible implementation, memory 22 may exist independently of processor 21 and may be connected to processor 21 via bus 24 to store instructions or program code. When processor 21 calls and executes the instructions or program code stored in memory 22, the intelligent text recognition method for scanned images provided in the following embodiments of the present invention can be implemented.

[0076] In another possible implementation, the memory 22 may also be integrated with the processor 21 .

[0077] Communication interface 23 is used to connect the intelligent text recognition device for scanned images to other devices via a communication network, such as Ethernet, a radio access network (RAN), or a wireless local area network (WLAN). Communication interface 23 may include a receiving unit for receiving data and a transmitting unit for transmitting data.

[0078] The bus 24 may be an Industry Standard Architecture (ISA) bus 24, a Peripheral Component Interconnect (PCI) bus 24, or an Extended Industry Standard Architecture (EISA) bus 24. The bus 24 may be divided into an address bus 24, a data bus 24, a control bus 24, and the like. For ease of representation, Figure 2 Only one thick line is used in the figure, but this does not mean that there is only one bus 24 or only one type of bus 24.

[0079] It should be pointed out that Figure 2The structure shown in the figure does not constitute a limitation on the intelligent text recognition device for scanning images, except Figure 2 In addition to the components shown, the intelligent text recognition device for scanned images may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0080] The following describes an intelligent text recognition method for scanned images provided by an embodiment of the present invention in conjunction with the accompanying drawings.

[0081] like Figure 3 As shown, the intelligent text recognition method for scanned images provided by the embodiment of the present invention includes the following steps 301 to 303.

[0082] 301. In response to a scanning operation performed by a user using a point reading device, a target image is acquired.

[0083] Optionally, the point reading device includes a scanning head and an intelligent recognition application installed in the point reading device. While the user is learning with the point reading device, the scanning head of the point reading device scans a target area on the book, which includes the content to be learned. The intelligent recognition application in the point reading device responds to the user's scanning operation using the point reading device to capture a target image. Exemplarily, the point reading device may be a reading pen or a scanning pen, and the scanning head may be a camera.

[0084] 302. Perform image recognition processing on the target image to determine image content included in the target image.

[0085] The image content may include content obtained by a user scanning with a point reading device.

[0086] Optionally, the image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture. In an embodiment of the present invention, after the point reading device acquires a target image, the image content of the target image can be determined in the following two ways.

[0087] In a possible implementation, the point reading device may use a pre-stored image recognition model to process the target image and generate image content, thereby determining the image content included in the target image.

[0088] In another possible implementation, the point reading device generates an image processing request based on the target image and sends the image processing request to the server. The server can use a pre-stored image recognition model to process the target image and generate image content, thereby determining the image content included in the target image.

[0089] 303. Delete irrelevant content in the image content to obtain target content.

[0090] The target content is the content in the image content that the user wants to learn, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering conditions. For example, the irrelevant content may be content that the user does not want to pay attention to. The point reading device or server deletes the irrelevant content in the image content to obtain the target content.

[0091] Using the intelligent text recognition method for scanned images provided by an embodiment of the present invention, a point-reading device or server performs image recognition processing on a target image captured by the point-reading device, determines the image content of the target image, and deletes irrelevant content from the image content to obtain the target content to be learned by the user. This eliminates the need for the user to manually select the target content again; the target content can be directly obtained. Compared to related technologies that require users to frequently use point-reading devices to scan to obtain the desired content, the intelligent text recognition method for scanned images provided by an embodiment of the present invention can solve the problem of frequent operation and poor user experience when users use point-reading devices for learning.

[0092] Combine Figure 3 ,like Figure 4 As shown, taking the image content including text format content as an example, the above step 303 may specifically include the following steps 401-402.

[0093] 401. Preprocess the image content to obtain words included in the processed content.

[0094] Optionally, the point-reading device or server deletes redundant information from the image content to obtain valid content, and determines at least one segmentation point for the valid content. The segmentation point is used to separate adjacent words in the valid content. The point-reading device or server then segments the valid content based on the at least one segmentation point to obtain the words included in the processed content.

[0095] For example, taking an image containing multiple words arranged from left to right, where the leftmost word is the first word scanned by the point reading device, the point reading device or server may delete redundant information from the image content to obtain valid content, which may include: the point reading device or server may input the multiple words into a pre-stored preprocessing model, delete redundant information from the multiple words using the preprocessing model, and output the processed multiple words. The redundant information may be information that is irrelevant to each of the multiple words. For example, the redundant information may be spaces and punctuation before the first word in the multiple words, extra spaces between two adjacent words in the multiple words, or spaces after the last word.

[0096] For example, if the image content is ",abc ac.", the valid content after preprocessing can be "abc ac." This valid content includes two segmentation points, one before "abc" and the other between "abc" and "ac." Therefore, the valid content includes the two words "abc" and "ac."

[0097] Optionally, when the image content consists of multiple words arranged from left to right, and there is a space and a period between two adjacent words, the point reading device or server needs to delete the space between the two adjacent words and replace the period with a comma. This can prevent the point reading device or server from identifying the first of the two adjacent words as a word abbreviation during subsequent processing.

[0098] For example, when the image content is "abc.def", the effective content after preprocessing is "abc,def".

[0099] 402. Delete irrelevant words from the words included in the processed content to obtain target content.

[0100] Optionally, non-relevant words can be incomplete words or words within a complete word that the user does not want to focus on. The processed content may include at least one word. When the processed content includes multiple words, the point reading device or server needs to delete the non-relevant words from the words included in the processed content to obtain the target content. When the processed content includes only one word, the point reading device or server directly determines that word as the target content.

[0101] Combine Figure 4 ,like Figure 5 As shown, the above step 402 may specifically include the following steps 501-502.

[0102] 501. When the processed content includes multiple words, delete incomplete words from the multiple words to obtain at least one complete word.

[0103] Optionally, when the processed content includes multiple words, the point reading device or server can determine incomplete words among the multiple words based on a pre-stored word list and delete the incomplete words to obtain at least one complete word. In an embodiment of the present invention, the point reading device or server can determine the incomplete words in the following two situations.

[0104] In one embodiment, when one of the multiple words is a single-letter word, the reading device or server confirms whether the single letter is "a" or "i". If the reading device or server determines that the single letter is not "a" or "i", the word including the single letter is determined to be an incomplete word.

[0105] It should be understood that the "a" or "i" here is not case sensitive.

[0106] In another case, when a word includes two or three letters, the reading device or server searches a pre-stored word list to find out whether the word is a complete word. If the word is not in the word list, the reading device or server confirms that the word is an incomplete word.

[0107] 502. Determine target content based on at least one complete word.

[0108] Optionally, when at least one remaining complete word is a complete word, the point reading device or server may determine the target content based on the at least one complete word in the following two cases.

[0109] In one embodiment, if at least one complete word satisfies a preset rule, the point reading device or server determines that the target content includes the at least one complete word. The preset rule may include at least one of the following: sentence, phrase, or proper noun. A proper noun is a type of noun that represents a specific, unique person or thing (such as a person's name, a place name, a country name, or a landscape name), as opposed to a common noun. In other words, if the content consisting of the at least one complete word is a sentence, phrase, or proper noun, the point reading device or server directly determines that the target content includes the at least one complete word.

[0110] When at least one complete word includes two words or three words arranged from left to right, the point reading device or server can determine whether the at least one complete word complies with the preset rules according to the following three methods.

[0111] In one possible implementation, if the point reading device or server determines that at least one complete word ends with a narrow punctuation mark, then the content of the at least one complete word is determined to be a sentence, i.e., the at least one complete word meets the preset rules. Punctuation marks in the narrow sense refer to specific symbols that serve as punctuation marks, such as a period, question mark, or comma.

[0112] Exemplarily, when the at least one complete word is "how are you?", and the at least one complete word includes the narrow punctuation mark "?", then the point reading device or server determines that the content composed of the at least one complete word is a sentence.

[0113] In another possible implementation, if the point reading device or server finds the at least one complete word in a pre-stored phrase table, it determines that the at least one complete word is a phrase, i.e., the at least one complete word meets the preset rules, and performs restoration processing on each word included in the at least one complete word. The restoration processing includes restoring the singular and plural forms of nouns, as well as the past tense and participle forms of verbs.

[0114] Exemplarily, when at least one complete word is "Looked after", the point reading device or server determines that the content consisting of the at least one complete word is a phrase, and restores the at least one complete word to "Look after".

[0115] In another possible implementation, if the point reading device or server determines that at least one complete word is a name, place name, or Chinese pinyin, then the at least one complete word is determined to be a proper noun, that is, the at least one complete word meets the preset rules.

[0116] Exemplarily, when the at least one complete word is "Tom", the point reading device or server determines that the content consisting of the at least one complete word is a proper noun.

[0117] Optionally, when at least one complete word includes four or more words arranged from left to right, the point reading device or server directly determines that the at least one complete word is a sentence, that is, the at least one complete word is the target content.

[0118] Optionally, when at least one complete word includes only one character, the point reading device or the server directly determines the at least one complete word as the target content.

[0119] In another case, if at least one complete word does not meet the preset rules, the point reading device or server adopts a preset evaluation strategy to determine the evaluation parameters of each complete word in at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

[0120] Optionally, if at least one complete word does not meet the preset rules, the point reading device or server needs to calculate the evaluation parameters of each complete word in the at least one complete word according to a preset evaluation strategy.

[0121] Exemplarily, the evaluation parameter is a total evaluation score of a complete word, and the total evaluation score may include: the sum of a length score of a complete word and a position score of a complete word. The length of a complete word is the number of letters in the complete word. The position score of a complete word is related to the position of the complete word in at least one complete word. When the complete word is the first word, the position score is 4; when the complete word is the second word, the position score is 2; and when the complete word is the third word, the position score is 0.

[0122] For example, when at least one complete word is “English, I of”, the evaluation parameters corresponding to the three complete words are 11, 3, and 2 respectively.

[0123] Optionally, after determining at least one evaluation parameter, determining the target content based on the at least one evaluation parameter may include: determining a target evaluation parameter that satisfies a parameter rule within the at least one evaluation parameter. If the target evaluation parameter is a single parameter, determining the complete word corresponding to the target evaluation parameter as the target content. If the target evaluation parameter includes multiple parameters, obtaining multiple complete words corresponding to the target evaluation parameter and determining the first word scanned by the point-reading device among the multiple complete words as the target content.

[0124] Exemplarily, determining the target evaluation parameter that satisfies the parameter rule may include: the point reading device or server sorting the at least one evaluation parameter in descending order; if there is only one maximum value in the at least one evaluation parameter, then the evaluation parameter with the largest value is the target evaluation parameter. If there are two maximum values ​​in the at least one evaluation parameter, then both evaluation parameters are the target evaluation parameters. If the difference between the two evaluation parameters ranked higher in the at least one evaluation parameter is less than a preset value, then both evaluation parameters are the target evaluation parameters.

[0125] For example, taking the preset value of 2 as an example. When at least one complete word is "English, I of", and the evaluation parameters corresponding to the three complete words are 11, 3, and 2 respectively, the target evaluation parameter is 11, and the target content is "English". When at least one complete word is "or greatjob", and the evaluation parameters corresponding to the three complete words are 6, 7, and 3 respectively, the target evaluation parameter is 7, and the target content is "great". When at least one complete word is "eat apple", and the evaluation parameters corresponding to the two complete words are 7 and 7 respectively, the target evaluation parameter is 7, and the target content is "eat".

[0126] After the point-reading device or server determines the target content, it needs to process the target content to determine the relevant information corresponding to the target content. The relevant information corresponding to the target content may include at least one of the following: part of speech, meaning, phonetic symbols, example sentences, voice data, and translation.

[0127] Optionally, the point-reading device or server includes a pre-stored translation model and speech synthesis model, wherein the translation model may include a dictionary translation model and a sentence translation model. The image processing request in step 301 may also include a target processing mode, which may be a translation mode and / or a speech conversion mode. The point-reading device or server selects translation and / or speech playback for the target content based on the target processing mode.

[0128] For example, the target processing mode includes a translation mode and a speech conversion mode. When the point reading device or server determines that the target content is a word, the word is input into the dictionary translation model, and the word's corresponding part of speech, meaning, phonetic symbol, and example sentence are output. Then, the point reading device or server inputs the word into the speech synthesis model, and outputs the speech data corresponding to the word. The user can learn the word based on the word's corresponding part of speech, meaning, phonetic symbol, example sentence, and speech data. Figure 6 As shown, when the target content is a word, the point reading device can display a word display interface, which includes: the word, and the word's corresponding part of speech, meaning, phonetic symbol and example sentence. The word display interface also includes a voice play button 61 and a recording button 62. When the user clicks the voice play button 61, the point reading device can play the pronunciation data of the word in response to the click operation. Based on the voice data, the user can perform listening training. When the user clicks the recording button 62, the point reading device can record the user's follow-up data in response to the click operation and obtain the user's follow-up voice data. After obtaining the user's follow-up voice data, the point reading device can obtain the evaluation result of the follow-up voice data according to a preset evaluation strategy. It should be understood that the evaluation strategy here is different from the evaluation strategy in step 502. The evaluation strategy here is used to evaluate the voice follow-up data and determine whether the user's pronunciation of the word is standard. As shown in FIG. Figure 6 As shown, the word display interface also includes a Return to Sentence button. When the point reading device or server determines that the target content is a word, but the user actually wants the target content to be a sentence, the user can click the Return to Sentence button. When the user clicks the Return to Sentence button, the point reading device or server can redefine the target content in response to the click operation.

[0129] For example, the target processing mode includes a translation mode and a speech conversion mode. When the point reading device or server determines that the target content is a sentence or phrase, the sentence or phrase is input into the sentence translation model, and the translation corresponding to the sentence or phrase is output. Then, the point reading device or server inputs the sentence or phrase into the speech synthesis model, and outputs the speech data corresponding to the sentence or phrase. The user can learn the sentence or phrase based on the translation and speech data corresponding to the sentence or phrase. Figure 7As shown, when the target content is a sentence, the point-reading device can display a sentence display interface. The sentence display interface includes: a sentence, phrase, or proper noun, and the corresponding translation of the sentence, phrase, or proper noun. The sentence display interface also includes a sentence audio playback button 71, a translation audio playback button 72, and a record button 73. When the user clicks the sentence audio playback button 71, the point-reading device can play the sentence's pronunciation data in response to the click. When the user clicks the translation audio playback button 72, the point-reading device can play the pronunciation data of the corresponding translation in response to the click. Based on the sentence audio data, the user can perform listening training and speaking practice. When the user clicks the record button 73, the point-reading device can record the user's follow-up reading data in response to the click, obtaining the user's follow-up reading audio data. After obtaining the user's follow-up reading audio data, the point-reading device can obtain an evaluation result for the follow-up reading audio data based on a preset evaluation strategy. It should be understood that the evaluation strategy here is different from the evaluation strategy in step 502. The evaluation strategy here is used to evaluate the voice follow-up data and determine whether the user's pronunciation of the sentence is standard.

[0130] It should be understood that if the server is the entity responsible for determining the target content and the related information corresponding to the target content, the server, after determining the related information corresponding to the target content, needs to send the target content and the related information corresponding to the target content to the point reading device. After the point reading device obtains the target content and the related information corresponding to the target content, it can display them on the display screen of the point reading device.

[0131] The above mainly introduces the solution provided by the embodiment of the present invention from the perspective of the intelligent text recognition device for scanned images. It can be understood that in order to realize the above functions, the intelligent text recognition device for scanned images includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0132] Figure 8 A possible schematic diagram of the composition of the intelligent text recognition device 800 for scanned images involved in the above embodiment is shown. Figure 8 As shown, the intelligent text recognition device 800 for scanned images may include: an acquisition unit 801 and a processing unit 802.

[0133] The acquisition unit 801 is configured to acquire a target image in response to a scanning operation performed by a user using a point reading device;

[0134] The processing unit 802 is used to perform image recognition processing on the target image obtained by the acquisition unit 801, determine the image content included in the target image, which includes the content obtained after the user uses the point reading device to scan, and delete the irrelevant content in the image content to obtain the target content, which is the content to be learned by the user in the image content, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering conditions.

[0135] Optionally, the image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture.

[0136] Optionally, when the image content includes content in text format, the processing unit 802 is specifically configured to:

[0137] The image content is preprocessed to obtain words included in the processed content, and irrelevant words in the words included in the processed content are deleted to obtain the target content.

[0138] Optionally, the processing unit 802 is specifically configured to:

[0139] Redundant information of the image content is deleted to obtain valid content, and at least one segmentation point of the valid content is determined, where the segmentation point is used to segment adjacent words in the valid content. The valid content is segmented based on the at least one segmentation point to obtain words included in the processed content.

[0140] Optionally, the processing unit 802 is specifically configured to:

[0141] When the processed content includes multiple words, incomplete words among the multiple words are deleted to obtain at least one complete word, and the target content is determined based on the at least one complete word.

[0142] Optionally, the processing unit 802 is specifically configured to:

[0143] If at least one complete word satisfies the preset rules, it is determined that the target content includes at least one complete word, and the preset rules include at least one of the following: sentence, phrase, proper noun; if at least one complete word does not satisfy the preset rules, a preset evaluation strategy is adopted to determine the evaluation parameters of each complete word in the at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

[0144] Optionally, the processing unit 802 is specifically configured to:

[0145] Among at least one evaluation parameter, determine the target evaluation parameter that meets the parameter rules; if the target evaluation parameter is one parameter, determine the complete word corresponding to the target evaluation parameter as the target content; if the target evaluation parameter includes multiple parameters, obtain multiple complete words corresponding to the target evaluation parameter, and determine the word that is first scanned by the point reading device among the multiple complete words as the target content.

[0146] Of course, the intelligent text recognition device 800 for scanned images provided by the embodiment of the present invention includes but is not limited to the above modules.

[0147] In actual implementation, the acquisition unit 801 and the processing unit 802 can be composed of Figure 2 The processor 21 shown calls the program code in the memory 22 to implement the process. Figures 3 to 5 The description of the intelligent text recognition method for scanned images shown in the figure will not be repeated here.

[0148] Another embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on the intelligent text recognition device 800 for scanning images, the intelligent text recognition device 800 for scanning images executes each step executed by the intelligent text recognition device 800 for scanning images in the method flow shown in the above method embodiment.

[0149] Another embodiment of the present invention provides a chip system, which is applied to an intelligent text recognition device 800 for scanning images. The chip system includes one or more interface circuits and one or more processors 21. The interface circuits and processors 21 are interconnected via circuits. The interface circuits are configured to receive signals from the memory 22 of the intelligent text recognition device 800 for scanning images and send the signals to the processor 21. The signals include computer instructions stored in the memory 22. When the processor 21 executes the computer instructions, the intelligent text recognition device 800 for scanning images performs the steps described in the method flow of the method embodiment.

[0150] In another embodiment of the present invention, a computer program product is also provided. The computer program product includes instructions. When the instructions are run on the intelligent text recognition device 800 for scanned images, the intelligent text recognition device 800 for scanned images executes the various steps executed by the intelligent text recognition device 800 for scanned images in the method flow shown in the above method embodiment.

[0151] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer-executable instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0152] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention shall be covered by the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for intelligent text recognition of scanned images, characterized in that: include: In response to a scanning operation of a user using a point reading device, acquiring a target image; Performing image recognition processing on the target image to determine image content included in the target image, wherein the image content includes content obtained after a user scans the target image using the point reading device; Deleting irrelevant content in the image content to obtain target content, where the target content is the content to be learned by the user in the image content, and the irrelevant content is the content in the image content that meets the user's negative feedback behavior triggering condition; When the image content includes text content, deleting irrelevant content in the image content to obtain target content includes: Preprocessing the image content to obtain words included in the processed content; Deleting irrelevant words from the words included in the processed content to obtain the target content; the irrelevant words are incomplete words or words in the complete words that the user does not want to pay attention to; The deleting of irrelevant words from the words included in the processed content to obtain the target content includes: When the processed content includes multiple words, deleting incomplete words from the multiple words to obtain at least one complete word; determining the target content according to the at least one complete word; The determining the target content according to the at least one complete word includes: If the at least one complete word satisfies a preset rule, determining that the target content includes the at least one complete word, the preset rule including at least one of the following: a sentence, a phrase, or a proper noun; If the at least one complete word does not meet the preset rule, a preset evaluation strategy is adopted to determine the evaluation parameters of each complete word in the at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

2. The intelligent text recognition method for scanned images according to claim 1, characterized in that: The image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture.

3. The intelligent text recognition method for scanned images according to claim 1, characterized in that: The pre-processing of the image content to obtain the words included in the processed content includes: Deleting redundant information of the image content to obtain valid content; determining at least one segmentation point of the valid content, where the segmentation point is used to segment adjacent words in the valid content; The valid content is segmented based on the at least one segmentation point to obtain words included in the processed content.

4. The intelligent text recognition method for scanned images according to claim 3, characterized in that: According to the At least one evaluation parameter, determining the target content, includes: Determining a target evaluation parameter that satisfies a parameter rule among the at least one evaluation parameter; If the target evaluation parameter is one parameter, determining the complete word corresponding to the target evaluation parameter as the target content; If the target evaluation parameter includes multiple parameters, multiple complete words corresponding to the target evaluation parameter are obtained, and the word scanned first by the point reading device among the multiple complete words is determined as the target content.

5. An intelligent text recognition device for scanned images, characterized in that: include: an acquisition unit, configured to acquire a target image in response to a scanning operation of a user using a point reading device; a processing unit configured to perform image recognition processing on the target image acquired by the acquisition unit, determine image content included in the target image, wherein the image content includes content obtained by a user after scanning using the point reading device, and delete irrelevant content from the image content to obtain target content, wherein the target content is content to be learned by the user in the image content, and the irrelevant content is content in the image content that meets a negative feedback behavior triggering condition of the user; When the image content includes content in text format, the processing unit is specifically configured to: Preprocessing the image content to obtain words included in the processed content; Deleting irrelevant words from the words included in the processed content to obtain the target content; the irrelevant words are incomplete words or words in the complete words that the user does not want to pay attention to; The processing unit is specifically configured to: When the processed content includes multiple words, deleting incomplete words from the multiple words to obtain at least one complete word; determining the target content according to the at least one complete word; The processing unit is specifically configured to: If the at least one complete word satisfies a preset rule, determining that the target content includes the at least one complete word, the preset rule including at least one of the following: a sentence, a phrase, or a proper noun; If the at least one complete word does not meet the preset rule, a preset evaluation strategy is adopted to determine the evaluation parameters of each complete word in the at least one complete word, obtain at least one evaluation parameter, and determine the target content based on the at least one evaluation parameter.

6. The intelligent text recognition device for scanned images according to claim 5, characterized in that: The image content is at least one of the following: a word, a phrase, a sentence, an article, or a picture.

7. The intelligent text recognition device for scanned images according to claim 5, characterized in that: The processing unit is specifically configured to: Deleting redundant information of the image content to obtain valid content; determining at least one segmentation point of the valid content, where the segmentation point is used to segment adjacent words in the valid content; The valid content is segmented based on the at least one segmentation point to obtain words included in the processed content.

8. The intelligent text recognition device for scanned images according to claim 7, characterized in that: The processing unit is specifically configured to: Determining a target evaluation parameter that satisfies a parameter rule among the at least one evaluation parameter; If the target evaluation parameter is one parameter, determining the complete word corresponding to the target evaluation parameter as the target content; If the target evaluation parameter includes multiple parameters, multiple complete words corresponding to the target evaluation parameter are obtained, and the word scanned first by the point reading device among the multiple complete words is determined as the target content.

9. An intelligent text recognition device for scanned images, characterized in that: The intelligent text recognition device for scanned images includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the intelligent text recognition device for scanned images executes the intelligent text recognition method for scanned images as described in any one of claims 1-4.

10. A computer-readable storage medium, characterized in that It comprises computer instructions, which, when executed on an intelligent text recognition device for scanned images, enable the intelligent text recognition device for scanned images to execute the intelligent text recognition method for scanned images as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Click-to-read data production method and system, storage medium and electronic equipment

    CN110490182A