A learning assistance method and system based on visual recognition
The visual recognition-based learning aid addresses slow and inaccurate scanning issues by integrating audio feedback and dynamic word management, enhancing learning efficiency and accuracy.
Patent Information
- Application Number
- CN202111496137.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing recognition devices have slow recognition speed, cumbersome operations, easy to misjudgment, and cannot accurately judge key words and sentences, resulting in low scanning and retrieval efficiency and students need to repeat operations.
Using a learning assistive method based on visual recognition, text information is obtained through the camera module, audio information is obtained by combining the voice recognition module, text and audio comparison and correction are carried out, and the meaning of the words to be checked is identified and displayed. Audio information is used to judge students' pronunciation accuracy, high-frequency simple words and review vocabulary database are eliminated, and scanning areas are optimized to improve recognition accuracy.
It realizes rapid and accurate identification and display of the meaning of the words to be looked up, improves scanning efficiency, saves learning time, strengthens students' spelling and reading aloud ability, and improves the accuracy of translation.
Smart Images

Figure CN114220110B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of learning assistance technology, and in particular to a learning assistance method and system based on visual recognition. Background Art
[0002] With the development of recognition technology, electronic dictionaries that rely on manual input have initially withdrawn from the market; a large number of auxiliary learning software that can search for questions by taking photos have appeared on the market. There are also multifunctional dictionary pens that can recognize the semantics and pronunciation of English or Chinese by swiping the dictionary pen on the text, and can also read aloud with the speaker. However, the existing recognition devices have a slow recognition speed, cumbersome operation process, and cannot perform other interactions during recognition.
[0003] The scanning width of the existing dictionary pen is constant. When using it to scan text, if the dictionary pen scans other text images at the same time, the dictionary pen will make a misjudgment, which will reduce the query efficiency. In this case, students can only manually select or cover other content and scan again. The scanning retrieval accuracy is low, and students may need to repeat the operation multiple times. When there is a lot of content to scan, the scanning process is long, and the dictionary pen cannot accurately judge the key words and sentences, and then they need to select them manually. Summary of the invention
[0004] In view of the above problems, the present invention aims to provide a learning assistance method and system based on visual recognition, which has accurate and fast recognition and can motivate students to read aloud.
[0005] To achieve the technical purpose, the solution of the present invention is: a learning assistance method based on visual recognition, the method comprising:
[0006] Students manually select the input mode, scan and input the text in the designated area of the book or homework, and recognize and output the text according to the image information to obtain the corresponding recognized text;
[0007] The students read aloud or follow the reading while scanning, collect the students' audio information, identify and extract the students' audio information, and obtain the corresponding reading text;
[0008] Compare the recognized text with the read text for correction, and display the corrected recognized text.
[0009] Preferably, the step of identifying and extracting the audio information of the student further includes:
[0010] Specify the pronunciation information and pause time of the text;
[0011] According to the students' initial settings, high-frequency, single-meaning simple words in the text to be read aloud are eliminated, and the recognized text is used as the initial text for comparison to mark the differences in the text to be read aloud, and the similarities and differences in the pronunciation and pronunciation information of the different words are analyzed.
[0012] Preferably, the recognized text includes adjacent words and words to be queried. When the student only reads out the words to be queried during reading, the recognized text is compared with the read text for correction, and the adjacent words are excluded to only display the meanings of the words to be queried.
[0013] Preferably, when the student cannot accurately pronounce the words to be queried, the whole sentence or part of the sentence is read while skipping the words to be queried. The recognized text is compared with the read text for correction, and then only the meanings of the words to be queried that are not read in the recognized text are displayed. When the read text is a complete sentence, the best meaning of the word to be queried in the sentence is preferably displayed first.
[0014] Preferably, when the pronunciation pause time of a certain word in the read text exceeds the threshold, the word is classified into the review word library;
[0015] The words to be queried are also classified into the review word library after being displayed.
[0016] Preferably, when the adjacent words belong to the high-frequency single-meaning simple words in the student's initial settings and do not belong to phrases, they are preferably excluded actively without being displayed;
[0017] When the adjacent words belong to the words in the review word library and the number of queries is less than the threshold, they are placed behind the words to be queried as alternative viewing content;
[0018] When the query word and the adjacent words form a phrase, the meaning of the phrase is preferably displayed first.
[0019] Preferably, when the read text or the scanned document is a whole sentence, the different sentences or the meanings of the words to be queried are preferably displayed first, and the translation of the whole sentence is used as alternative viewing content.
[0020] Preferably, when the words to be queried are in the same sentence and not adjacent, and there are multiple meaning results for the words to be queried, the best meaning of the above words to be queried in the sentence is preferably displayed first, and a switching control is also displayed. The student can view other meanings of the words to be queried by swiping.
[0021] A learning assistance system adopts a learning assistance method based on visual recognition, including:
[0022] A camera module, which can scan and obtain the text information of a specified scanning area, and the text range of the scanning area is larger than the range of the words to be queried;
[0023] A voice recognition module, which can obtain the audio information of the student;
[0024] An audio processing module, which can recognize the read text, pronunciation information and pause time corresponding to the audio information;
[0025] A comparison and analysis module compares the recognized text with the recited text for correction, and displays the corrected recognized text.
[0026] A display module displays the meaning or pronunciation of the words and sentences.
[0027] It is configured to query the text to be queried to obtain a corresponding query result and display the query result.
[0028] The beneficial effects of the present invention are as follows: The method of this application can obtain the scanned text information and audio information. Through the comparison of the two, the word to be searched that the student needs to view can be quickly locked. The recognition speed is fast and the accuracy is high, which can save learning time. At the same time, when reading the whole sentence aloud, the word to be searched is scanned locally, and the word to be searched can be better translated in combination with the context, and the given meaning is also closer to the true meaning in the sentence, and the translation is more accurate. The method of this application can effectively strengthen the student's word spelling ability and sentence reading ability, and can more quickly find the words with inaccurate pronunciation and unfamiliar words. Description of the Drawings
[0029] Figure 1 It is a schematic structural diagram of Embodiment 1 of the present invention;
[0030] Figure 2 It is a schematic structural diagram of Embodiment 2 of the present invention;
[0031] Figure 3 It is a schematic structural diagram of Embodiment 3 of the present invention;
[0032] Figure 4 It is a schematic structural diagram of Embodiment 4 of the present invention. Detailed Description of the Invention
[0033] The present invention will be further described in detail below with reference to the drawings and specific embodiments. The specific embodiments listed below are exemplary rather than restrictive. The terms "including" and "having" and their conventional variant expressions in the following specific embodiments are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. Any minor modifications, equivalent replacements and improvements made to the following specific embodiments based on the technical essence of this application shall be included in the protection scope of the technical solution of this application.
[0034] Embodiment 1:
[0035] Such as Figure 1As shown, the present invention discloses a learning assistance method based on visual recognition. When scanning text, students can actively read sentences according to their experience, actively try to spell unfamiliar words that can be spelled, and skip words that cannot be spelled. This method can lock the words to be queried faster and at the same time allow students to actively participate and check whether their pronunciation is accurate. The method includes:
[0036] S101. According to the initial settings of the students, eliminate simple words with high frequency and single meaning in the read text (definite articles, simple nouns, such as the, a, if, that, apple). The students manually select the input mode as the reading mode (the whole sentence can be scanned or part of the sentence can be scanned), scan and input the text in the designated area of the book, and identify and output according to the image information of the text to obtain the corresponding recognized text;
[0037] S102. While scanning, the students read the text, collect the audio information of the students, identify and extract the audio information of the students to obtain the corresponding read text, the pronunciation information of the designated text, and the pause time;
[0038] S103. If the students do not read the whole sentence or part of the sentence and skip the words to be queried, compare and correct the recognized text with the read text, and only display the meanings of the words to be queried that are not read in the recognized text. When the read text is a complete sentence, the best meaning of the word to be queried in the sentence is preferentially displayed;
[0039] S104. If the pronunciation pause time of a certain word in the read text exceeds the threshold, classify the word into the review word library; the words to be queried are also classified into the review word library after being displayed.
[0040] First, students who use the translation pen alone usually have a certain foundation in language grammar and can recognize some words. The purpose of using the translation pen is to solve the problem of unclear meanings of some words and sentences. Moreover, many students have also learned natural phonics. If they encounter an unfamiliar word, they can also read out its approximate pronunciation according to the spelling characteristics. Since the scanning speed of the translation pen is slow, when the translation pen sweeps across the specified text, students can actively read aloud or follow along. By pronouncing the words, they can also strengthen their understanding of sentences or words and make more efficient use of their study time (while also honing their phonetic skills). Take "If the dream is big enough, the facts don't count" as an example. In the reading mode, as the translation pen sweeps across the text, the student reads along. The student doesn't know the words "dream" and "enough", but can roughly pronounce "dream", and can't pronounce "enough". When the translation pen recognizes the pronunciation of "dream", there is an obvious pause. When "enough" is scanned but not read out, it is determined that the student has not mastered these two words yet. Then the meanings of these two words are displayed in the middle, and the translation of the whole sentence is provided as an alternative for viewing. At the same time, these two words are included in the review word bank.
[0041] Embodiment 2:
[0042] The present invention discloses a learning assistance method based on visual recognition. As Figure 2 shown, through following along, students can make more efficient use of the scanning time. At the same time, the following-along difficulty is low and it is easier to accept. Through oral training, they can also better familiarize themselves with the words to be looked up. The method includes:
[0043] S201. The student manually selects the input mode as the following-along mode, usually for scanning the whole sentence. The translation pen first starts to scan the text in the specified area (using punctuation marks as the segmentation area), and the translation pen plays the recognized text, and the student follows along.
[0044] S202. The translation pen synchronously collects the student's audio information, recognizes and extracts the student's audio information to obtain the corresponding read text, the pronunciation information of the specified text, and the pause time.
[0045] S203. The student eliminates the simple words with high-frequency single meanings in the read text, uses the recognized text as the initial text to compare and mark the differences in the read text, and analyzes the similarities and differences between the pronunciations of the different words and the pronunciation information.
[0046] S204. When the pronunciation pause time of a certain word in the read text exceeds the threshold, the word is included in the review word bank; the words to be looked up are also included in the review word bank after being displayed.
[0047] In the follow-up reading mode, students can follow the voice played by the translation pen and read along. According to the accuracy of word pronunciation and pause time, the students' proficiency in the word or phrase can be judged, and the words to be checked and the words that are not pronounced standardly can be quickly identified. Moreover, following along can reduce the difficulty of students' participation. Reading along can also strengthen the understanding of words and sentences and improve learning effects.
[0048] Embodiment three:
[0049] The present invention discloses a learning assistance method based on visual recognition, such as Figure 3 As shown, in order not to affect the entire reading process, local scanning and recognition can also be performed efficiently to quickly understand the meaning of the word in the sentence. The method includes:
[0050] S301, the student manually selects the input mode as the reading mode, and the student reads a sentence. At the same time, the translation pen collects the student's audio information, recognizes and extracts the student's audio information, and obtains the corresponding reading text;
[0051] S302, when encountering a word to be searched, stop reading aloud or try to spell it, and scan and enter the word to be searched at the same time (since the scanning area may be larger than the word to be searched, adjacent words will be recognized), and recognize and output the image information of the text to obtain the corresponding recognized text;
[0052] S303, when the recognized text contains adjacent words and the word to be searched, the whole sentence is read aloud while skipping the word to be searched or the pronunciation is inaccurate (or the pause time for spelling is too long), the recognized text is compared and corrected with the read text, and the adjacent words are removed to display only the meaning of the word to be searched;
[0053] S304, displaying the best meaning of the above-mentioned word to be searched in the sentence, and displaying a switching control at the same time, so that students can view other meanings of the word to be searched by sliding.
[0054] Usually, the reading speed is slightly faster than the scanning speed of the translation pen when students read a whole sentence or chapter. Start the reading mode of the translation pen, and students read normally. If they encounter a word that they don't understand the meaning or spelling, they can skip it and use the translation pen to scan the word to be checked. After reading aloud, check the meaning of the word.
[0055] First, the translation pen continuously records the student's audio information, and can translate the words to be searched through the context or the words before and after, so that the translation of the words to be searched is more accurate. It does not affect the speed of reading aloud, and only needs to wait for a few seconds when scanning the words. The method of this application does not require scanning the entire sentence, and the input efficiency is high. Students also strengthen their understanding of the sentence through reading aloud.
[0056] Embodiment 4:
[0057] The present invention discloses a learning assistance method based on visual recognition, as Figure 4 shown, the method includes:
[0058] S401. The student manually selects the input mode as the reading mode, and the student reads the sentence. Meanwhile, the translation pen collects the student's audio information, recognizes and extracts the student's audio information, and obtains the corresponding read text.
[0059] S402. The recognized text includes adjacent words and the word to be queried. When the student only reads the word to be queried during reading, but the pronunciation of the word to be queried is inaccurate or the spelling time exceeds the threshold, the meaning of the word to be queried is displayed.
[0060] When the student only reads the adjacent word during reading, but the pronunciation of the adjacent word is accurate and there is no pause, the meaning of the word to be queried is displayed.
[0061] If the adjacent word and the word to be queried form a phrase, the meaning of the phrase is displayed.
[0062] For example, for "If the dream is big enough, the facts don't count", when the student wants to query the meaning of enough (big is the adjacent word), since big enough forms a phrase, the meaning of big enough is displayed. When the student wants to query the meaning of dream (the and is are adjacent words), since both the and is are simple words with high-frequency single meanings, the meaning of dream is displayed.
[0063] A learning assistance system that adopts the learning assistance method based on visual recognition, includes: a camera module that can scan and obtain the text information of a specified scanning area, and the text range of this scanning area is larger than the range of the word to be queried; a voice recognition module that can obtain the student's audio information; an audio processing module that can recognize the read text, pronunciation information and pause time corresponding to the audio information; a comparison and analysis module that corrects by comparing the recognized text with the read text and displays the corrected recognized text; a display module that displays the meaning or pronunciation of the sentence. It is configured to query the text to be queried to obtain the corresponding query result and display the query result.
[0064] In the specific embodiments of the present application, the magnitudes of the sequence numbers of each process do not necessarily mean the inevitable sequence of execution. The execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0065] In each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0066] When the above-mentioned functional unit is implemented in software form and sold or used as an independent product, it may be stored in a memory accessible by a computer device. Therefore, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device to execute some or all of the steps of the above-mentioned methods in each embodiment of the present application.
Claims
1. A learning assistance method based on visual recognition, characterized in that, The method includes: The student manually selects an input mode, scans and enters the text in the specified area of the book or assignment, and recognizes and outputs according to the image information of the text to obtain the corresponding recognized text. While scanning, the student reads or follows the text aloud, collects the student's audio information, recognizes and extracts the student's audio information to obtain the corresponding read text. Compare the recognized text with the read text for correction and display the corrected recognized text. In the recognition and extraction of the student's audio information, it also includes: The pronunciation information and pause time of the specified text. According to the student's initial setting, eliminate the simple words with high-frequency single meanings in the read text, use the recognized text as the initial text to compare and mark the differences in the read text, and analyze the similarities and differences in the pronunciation of the different words and the pronunciation information. The recognized text includes adjacent words and words to be queried. When the student only reads out the words to be queried during reading, compare and correct the recognized text with the read text, eliminate the adjacent words and only display the meanings of the words to be queried. When the student cannot accurately pronounce the words to be queried, read the whole sentence or part of the sentence and skip the words to be queried, compare and correct the recognized text with the read text, and only display the meanings of the words to be queried that are not read in the recognized text. When the read text is a complete sentence, the best meaning of the word to be queried in the sentence is preferably displayed.
2. The learning assistance method based on visual recognition according to claim 1, wherein: When the pronunciation pause time of a certain word in the read text exceeds the threshold, the word is classified into the review word bank. The words to be queried are also classified into the review word bank after being displayed.
3. The learning assistance method based on visual recognition according to claim 2, characterized in that: When the adjacent words belong to the simple words with high-frequency single meanings in the student's initial setting and do not belong to phrases, they are preferably eliminated actively and not displayed. When the adjacent words belong to the words in the review word bank and the number of queries is less than the threshold, they are placed behind the words to be queried as alternative viewing content. When the query word and the adjacent word form a phrase, the meaning of the phrase is preferably displayed.
4. The learning assistance method based on visual recognition according to claim 2, wherein: When the read text or the scanned file is a whole sentence, the meanings of the different sentences or the words to be queried are preferably displayed, and the translation of the whole sentence is used as alternative viewing content.
5. The learning assistance method based on visual recognition according to claim 1, wherein: When the words to be queried are in the same sentence and not adjacent, and there are multiple meaning results for the words to be queried, the best meaning of the above words to be queried in the sentence is preferably displayed, and a switching control is also displayed. The student can view other meanings of the words to be queried by swiping.
Citation Information
Patent Citations
Oral pronunciation correction method and device, equipment and storage medium
CN113393864A