Relay reading training method and server for executing the same
The relay reading education method and server enhance reading habits by engaging users in active reading with real-time error detection and gamified feedback, addressing the limitations of passive e-book platforms in improving reading concentration and comprehension.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- CAREERABLE CO LTD
- Filing Date
- 2025-08-18
- Publication Date
- 2026-07-29
AI Technical Summary
Existing e-book platforms fail to effectively encourage users, particularly elementary school students and the elderly, to develop a habit of thorough reading by actively participating in e-book content, as they primarily focus on passive text viewing or audio reading, which does not significantly improve reading concentration or comprehension.
A relay reading education method and server that facilitate a reading group setting where users read aloud, provide real-time feedback on errors, and switch reading authority based on performance, using AI-driven speech recognition and gamified visual feedback to enhance engagement and accuracy.
The method and server support learners in developing a close reading habit by actively participating in reading activities, detecting errors, and providing interactive learning environments that promote continuous engagement and improved reading skills through gamified feedback.
Smart Images

Figure 112025093721441-PAT00010_ABST
Abstract
Description
Technology Field
[0001] Embodiments of the present invention relate to a relay reading education method and a server for performing the same, and more specifically, to a relay reading education method and a server for performing the same that supports a learner in effectively developing a close reading habit. Background Technology
[0002] Due to the recent excessive consumption of short-form content, an increasing number of people have become accustomed to brief and stimulating information. Consequently, there is a growing number of cases where people struggle to concentrate on reading long texts to the end, or complain of declining reading comprehension and difficulty sustaining their studies. This phenomenon is emerging as a social issue that goes beyond mere individual reading habits, causing a large number of people to experience a decline in their overall understanding of texts and their ability to sustain learning.
[0003] However, most existing e-book platforms focus solely on text viewing functions and have failed to provide an environment where people can actively participate in e-book content. Although some e-book platforms support e-book reading services using speech synthesis technology, there are limitations in encouraging people to develop a habit of thorough reading.
[0004] In particular, for elementary school students whose reading comprehension skills are not yet fully developed, or for the elderly who need to maintain cognitive function through repetitive learning, simply viewing text or listening to audio readings is not expected to have a substantial effect on improving reading concentration. For these individuals, a learning method is required that enables the formation of reading habits by having them read e-books aloud and repeatedly experience feedback on errors that occur during the process.
[0005] Therefore, it is necessary to introduce a reading aloud education method that can support learners in actively participating in reading aloud and naturally forming a habit of close reading through repetitive learning. Prior art literature
[0006] Korean Registered Patent No. 10-1566013 (Title of Invention: Method and System for Providing Improvement in Speech Accuracy and Expressiveness through Electronic Book Reading) The problem to be solved
[0007] The present invention aims to solve various problems, including those mentioned above, by providing a relay reading education method that supports learners in effectively developing a close reading habit, and a server for performing the same. However, these problems are exemplary and the scope of the present invention is not limited by them. means of solving the problem
[0008] A relay reading education method according to one embodiment of the present invention may include the steps of: receiving a request to open a reading group from a first terminal (the reading group includes a first terminal acting as a teacher and a second-1 terminal and a second-2 terminal acting as learners); providing an e-book text to the reading group; granting voice input rights to the second-1 terminal; receiving first reading data from the second-1 terminal and generating first error information including the number of reading errors and time limit exceeding information based on the first reading data; transmitting first visual information to the second-1 terminal to visually display the reading status based on the first error information; and granting voice input rights to the second-2 terminal or the first terminal based on the first error information.
[0009] The step of providing e-book text to the reading group may include receiving a request to provide e-book text from the first terminal, providing a pre-stored list of e-books to the first terminal, receiving information about an e-book selected from the list of e-books from the first terminal, and dividing the entire text constituting the selected e-book into sentence units, and then sequentially providing a pre-set number of sentences to the reading group.
[0010] The step of generating the first error information may include receiving the first reading data from the 2-1 terminal, applying the first reading data to a pre-trained speech recognition model to convert it into a first speech recognition text, and generating the first error information based on the first speech recognition text and the e-book text provided to the reading group.
[0011] The above first error information may include the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit.
[0012] The step of transmitting the first time information may include the step of generating user interface display information based on the first error information, the step of transmitting the user interface display information to the second-1 terminal, and the step of controlling the state of at least one of the graphic elements constituting the user interface displayed on the second-1 terminal to change according to the user interface display information, and the user interface display information may include at least one of i) information instructing that some of the pre-displayed multiple segments disappear as the number of errors increases, ii) information instructing that the display state of the pre-displayed time progress bar be changed according to the elapsed time, and iii) information instructing that the display form of the pre-displayed icon be changed according to the error level (the error level means a level calculated based on the number of errors and the elapsed time).
[0013] The above relay reading education method may further include the steps of: generating a first modified text by modifying an e-book text provided to the reading meeting; generating a first synthetic voice data by performing Text-to-Speech (TTS) on the first modified text; transmitting the first synthetic voice data to the first terminal and the second-1 terminal; generating second error information based on the first synthetic voice data; and granting voice input rights to the first terminal or the second-1 terminal based on the second error information.
[0014] The first modified text above may be a text generated by at least one of the following text generation methods: a method of randomly changing the e-book text provided to the reading meeting by a preset ratio, or a method of additionally inserting arbitrary characters into the e-book text provided to the reading meeting.
[0015] In a server that performs operations necessary for relay reading education according to an embodiment of the present invention, the server receives a request to open a reading group from a first terminal (the reading group includes a first terminal acting as a teacher and a second-1 terminal and a second-2 terminal acting as a learner), provides an e-book text to the reading group, grants voice input authority to the second-1 terminal, receives first reading data from the second-1 terminal, generates first error information including the number of reading errors and time limit exceeded information based on the first reading data, transmits first visual information to the second-1 terminal to visually display the reading status based on the first error information, and may grant voice input authority to the second-2 terminal or the first terminal based on the first error information.
[0016] The server receives a request to provide e-book text from the first terminal, provides a pre-stored list of e-books to the first terminal, receives information about an e-book selected from the list of e-books from the first terminal, divides the entire text constituting the selected e-book into sentences, and then sequentially provides a pre-set number of sentences to the reading group.
[0017] The server receives the first reading data from the second-1 terminal, applies the first reading data to a pre-trained speech recognition model to convert it into a first speech recognition text, and can generate the first error information based on the first speech recognition text and the e-book text provided to the reading group.
[0018] The above first error information may include the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit.
[0019] The server may generate user interface display information based on the first error information and transmit the user interface display information to the second-1 terminal, and control the state of at least one of the graphic elements constituting the user interface displayed on the second-1 terminal to be changed according to the user interface display information, and the user interface display information may include at least one of i) information instructing that some of the pre-displayed multiple segments be extinguished as the number of errors increases, ii) information instructing that the display state of the pre-displayed time progress bar be changed according to the elapsed time, and iii) information instructing that the display form of the pre-displayed icon be changed according to the error level (the error level means a level calculated based on the number of errors and the elapsed time).
[0020] The server may generate a first modified text by modifying the e-book text provided to the reading meeting, generate first synthetic voice data by performing text-to-speech (TTS) on the first modified text, transmit the first synthetic voice data to the first terminal and the second-1 terminal, generate second error information based on the first synthetic voice data, and grant voice input rights to the first terminal or the second-1 terminal based on the second error information.
[0021] The first modified text above may be a text generated by at least one of the following text generation methods: a method of randomly changing the e-book text provided to the reading meeting by a preset ratio, or a method of additionally inserting arbitrary characters into the e-book text provided to the reading meeting. Effects of the invention
[0022] According to one embodiment of the present invention as described above, a relay reading education method and a server for performing the same can be implemented to support learners in effectively developing a close reading habit. Of course, the scope of the present invention is not limited by such effects.
[0023] A relay reading education method according to one embodiment of the present invention can support learners in effectively forming a reading habit by inducing learners to continuously participate in reading activities.
[0024] In addition, the relay reading education method according to the present invention can detect reading errors by the learner and automatically switch the reading authority when a reading error is detected, thereby encouraging the learner to actively participate in reading activities while maintaining immersion.
[0025] The relay reading education method according to the present invention can provide an interactive learning environment that goes beyond existing passive learning methods by inducing active participation and learning motivation of actual learners through reading errors intentionally caused by a virtual user. Brief explanation of the drawing
[0026] FIG. 1 is a block diagram schematically showing the configuration of a relay reading education system according to one embodiment of the present invention. Figure 2 is a conceptual diagram showing the server of Figure 1. FIG. 3 is a flowchart schematically illustrating a relay reading education method according to one embodiment of the present invention. FIG. 4 is a flowchart illustrating in more detail the step (S120) of providing the e-book text of FIG. 3. FIG. 5 is a flowchart illustrating in more detail the step (S140) of generating the first error information of FIG. 3. FIG. 6 is a flowchart showing the step (S150) of transmitting the first time information of FIG. 3 in more detail. Figure 7 is a flowchart schematically illustrating a relay reading education method that includes the participation of artificial intelligence in reading. Figures 8 and 9 are exemplary drawings showing a user interface displayed during a relay reading meeting. FIG. 10 is a flowchart schematically illustrating a method for classifying reading error types according to an embodiment of the present invention. Figure 11 is a flowchart schematically illustrating a method for classifying reading error types, including providing feedback information. Specific details for implementing the invention
[0027] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0028] In describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description is omitted.
[0029] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0030] A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0031] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0032] The attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification, and the technical concept disclosed in this specification is not limited by the attached drawings; it should be understood that all modifications, equivalents, and substitutions included within the concept and technical scope of the present invention are included.
[0033] Throughout the specification, when a part is described as being “connected” to another part, this includes not only cases where they are “directly connected” but also cases where they are “electrically connected” with other elements interposed between them. Furthermore, when a part is described as “comprising” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0034] In this specification, the term “and / or” includes any and all combinations of one or more of the associated listed items. For example, the expression “A and / or B” indicates A, B, or A and B. Expressions such as “at least one” may be used to refer to one or more of a plurality of components. For example, expressions such as “at least one of a, b, and c” or “at least one selected from the group consisting of a, b, and c” may indicate “a”, “b”, “c”, “a, b”, “b, c”, “a, c”, or “a, b, c”.
[0035] In this specification, terms such as "substantially," "approximately," and similar terms are used as terms of approximation rather than terms of degree, and may be used to describe inherent variations in measured or calculated values that a person skilled in the art can recognize. For example, terms such as "may" or "may" may be used to mean "one or more embodiments disclosed in this specification."
[0036] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0038] Relay reading education method and a server for performing the same
[0040] FIG. 1 is a block diagram schematically showing the configuration of a relay reading education system according to one embodiment of the present invention.
[0041] As illustrated in FIG. 1, the relay reading education system according to the present invention may include a server (100), a first terminal (110), and at least one second terminal (120, 130…). For convenience of explanation, the first of the second terminals is referred to as the second-1 terminal (120) and the second as the second-2 terminal (130), but the number of second terminals is not limited thereto.
[0042] The first terminal (110), the second-1 terminal (120), and the second-2 terminal (130) may be connected to the server (100) via a wireless network (Wi-Fi, LTE, 5G, etc.) and configured to transmit and receive data in real time.
[0043] The first terminal (110), the second-1 terminal (120), and the second-2 terminal (130) may all be mobile terminals or fixed terminals. Mobile terminals include, but are not limited to, smartphones, tablet PCs, PDAs, laptops, etc. Fixed terminals may include desktops, etc.
[0044] The first terminal (110) can be used by the first user (e.g., teacher, instructor, facilitator, etc.) who hosts the reading meeting.
[0045] The first user can request the opening of the reading meeting to the server (100) through the first terminal (110).
[0046] The server (100) can receive a request to open the reading meeting from the first terminal (110).
[0047] The above reading session refers to a session configured for multiple terminals to sequentially read an e-book text.
[0048] An e-book refers to electronic content provided in formats such as PDF, EPUB, and TXT.
[0049] E-book text refers to sentence-unit text data extracted from e-books and used for reading aloud education.
[0050] The server (100) can receive a request to provide e-book text from the first terminal (110) and can provide a list of e-books that are previously stored to the first terminal (110).
[0051] The above list of e-books may include e-book identification numbers, e-book titles, and author information.
[0052] The first user can view the list of e-books received from the server (100) through the first terminal (110), select an e-book suitable for reading according to a predetermined e-book selection criterion, and then transmit information about the selected e-book to the server (100) through the first terminal (110).
[0053] The above e-book selection criteria may include the age group of the participants in the reading education, the reading level, and the learning objectives of the reading group.
[0054] After receiving information about the selected e-book from the first terminal (110), the server (100) can load the selected e-book based on the identification information of the selected e-book.
[0055] The server (100) can divide the entire text constituting the selected e-book into sentence units and then sequentially provide a preset number of sentences to the reading group.
[0056] When the server (100) divides text into sentence units, it may use a sentence separation method based on sentence termination symbols (e.g., periods, question marks, etc.) recognition or natural language processing.
[0057] For example, assuming that the selected e-book consists of a total of four sentences such as ‘I wake up early in the morning. I go to school with a friend. I listen attentively in class. I play soccer during lunchtime,’ and that the preset number is ‘2,’ the server (100) can divide the entire text constituting the e-book into four sentences such as ‘I wake up early in the morning.’, ‘I go to school with a friend.’, ‘I listen attentively in class.’, and ‘I play soccer during lunchtime,’ and then provide the first two sentences, ‘I wake up early in the morning. I go to school with a friend.’ to the reading group, and then provide the remaining sentences, ‘I listen attentively in class. I play soccer during lunchtime.’
[0058] The server (100) can grant voice input rights to the 2-1 terminal (120).
[0059] The aforementioned voice input authorization system can provide the following educational benefits. First, systematic management of speech order prevents confusion among learners and enables smooth class progress. Second, the automatic voice input authorization switching mechanism encourages continuous learning participation without teacher intervention. Third, authorization is immediately switched in the event of an error, increasing opportunities for active learner participation and mutual correction.
[0060] The 2-1 terminal (120) may be a terminal used by a second user, who is one of the learners participating in the reading group.
[0061] When the second user is granted voice input permission to the second-1 terminal (120), he / she activates the microphone pre-installed in the second-1 terminal (120) and then reads aloud the text of the e-book provided to the reading meeting.
[0062] The 2-1 terminal (120) can receive the voice of the 2nd user reading through a mounted microphone and generate the 1st reading data.
[0063] Reading data refers to audio data containing the content of an e-book text provided to the reading group read aloud by a learner, which can be stored in formats such as WAV, MP3, FLAC, etc.
[0064] The second user can transmit the first reading data to the server (100) through the second-1 terminal (120).
[0065] The server (100) receives the first reading data from the second-1 terminal (120) and can generate first error information including the number of reading errors and timeout information based on the first reading data.
[0066] The server (100) receives the first reading data from the second-1 terminal (120) and can convert the first reading data into a first speech recognition text by applying it to a pre-trained speech recognition model.
[0067] For example, the server (100) can input the first reading data into the speech recognition model to generate the first speech recognition text, “I wake up in the morning. I go to school with a friend.”
[0068] A speech recognition model refers to an artificial intelligence model that receives speech data as input and converts the utterances contained in the speech data into text format, and deep learning speech recognition models based on RNN (Recurrent Neural Network), CNN (Convolutional Neural Network), or Transformer can be utilized.
[0069] Speech recognition text refers to text data converted by inputting reading data into a speech recognition model.
[0070] The server (100) can generate the first error information based on the first voice recognition text and the e-book text provided to the reading group.
[0071] The above first error information may include the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit.
[0072] The above first error information may include the cumulative utterance time measured during the reading performance, the difference value from the above time limit, etc.
[0073] The above first error information may include reading error type information output by the reading error type classification method described below. A detailed description of the reading error type classification method will be provided later.
[0074] The server (100) can divide the first voice recognition text and the e-book text into words based on spacing, and then calculate the number of errors by comparing the degree of alignment between the divided words.
[0075] The server (100) can calculate the word matching degree based on a matching method based on a minimum edit distance (e.g., Levenshtein distance).
[0076] For example, the above e-book text is ‘I in the morning early I get up. With a friend. At school If the first voice recognition text is ‘I wake up in the morning. I go to school with a friend.’, the server (100) can calculate the number of errors by comparing the e-book text and the first voice recognition text word by word. It can be calculated as a total of ‘2’ errors, with ‘early’ being omitted in the first sentence and ‘school’ being incorrectly recognized as ‘school’ in the second sentence.
[0077] In addition, if the total utterance time of the first reading data is '15 seconds' and the preset time limit is '10 seconds', the server (100) can calculate the elapsed time as '5 seconds'.
[0078] The server (100) can transmit first visual information to the second-1 terminal (120) to visually display the reading status based on the first error information.
[0079] The visual feedback system described above provides effective learning motivation while minimizing the learner's cognitive load. By visually simplifying complex numerical information through intuitive graphic elements (heart segments, time progress bars, and emotion emoticons), learners can immediately recognize their learning status. This real-time visual feedback can enhance learning immersion through gamification effects and support the development of self-regulated learning abilities.
[0080] For example, the first visual information may refer to information for visually providing feedback on the learner's reading errors and the passage of time. The first visual information may be visual feedback data that allows the learner to intuitively recognize their reading status by changing the color, shape, display state, etc., of graphic elements (e.g., segments, progress bars, icons, etc.) on the user interface.
[0081] The server (100) generates user interface display information based on the first error information and transmits the user interface display information to the second-1 terminal (120), and can control the state of at least one of the graphic elements constituting the user interface displayed on the second-1 terminal (120) to change according to the user interface display information.
[0082] The above user interface display information may include at least one of i) information instructing that some of the pre-displayed multiple segments disappear as the number of errors increases, ii) information instructing that the display state of the pre-displayed time progress bar change according to the elapsed time, and iii) information instructing that the display form of the pre-displayed icon change according to the error level.
[0083] The error level refers to a level calculated based on the number of errors and elapsed time, and can be pre-set by the developer so that the error level increases as the number of errors and elapsed time increase.
[0084] The error level can be calculated by the following formula:
[0085] [Mathematical Formula 1]
[0086] Error Level = α × (Number of Errors / Maximum Allowable Errors) + β × (Elapsed Time / Timeout)
[0087] Here,
[0088] - α, β: weights (0 ≤ α, β ≤ 1, α + β = 1)
[0089] - Maximum allowed errors: Preset value (e.g., 5 times)
[0090] - Error level: Normalized value between 0 and 1
[0091] For example, α can be set to 0.6 and β to 0.4, and the calculated error level can be normalized to a value between 0 and 1.
[0092] As such, the relay reading education method according to the present invention can overcome the limitations of existing reading education that relied on the subjective judgment of teachers by establishing an objective evaluation system through mathematical modeling. In particular, the accuracy and fluency of reading can be evaluated simultaneously through the error level calculation formula (Mathematical Formula 1), thereby promoting balanced improvement in reading ability.
[0093] By adjusting the weights α and β in the above mathematical formula 1, the following effects can be achieved. Increasing the value of α (e.g., 0.8) enables an accuracy-centered evaluation, which is effective for correcting the pronunciation of beginner learners, while increasing the value of β (e.g., 0.8) enables a fluency-centered evaluation, which is effective for improving the reading speed of intermediate or higher-level learners. Through this adjustment of weights, customized evaluation tailored to the learner's level and learning goals is possible.
[0094] A graphic element composed of five heart-shaped segments may be included on the user interface displayed on the second-1 terminal (120). Based on the first error information, the server (100) may generate user interface display information instructing that one heart-shaped segment be destroyed each time the number of errors increases by one, and transmit this information to the second-1 terminal (120). Based on the received user interface display information, the second-1 terminal (120) may change the display state of the heart-shaped segments (e.g., the color of the heart-shaped segments, the shape of the heart-shaped segments, etc.) according to the number of errors.
[0095] A time progress bar indicating the elapsed time limit may be included on the user interface displayed on the second-1 terminal (120). Based on the first error information, the server (100) may generate information instructing the time progress bar to blink red when the second user's speech time approaches the time limit and transmit this information to the second-1 terminal (120). The second-1 terminal (120) may change the display state of the time progress bar based on the received user interface display information. If the second user's speech time exceeds the time limit, the server (100) may generate user interface information instructing the elapsed time to be displayed in a numerical form (e.g., '+5 seconds') based on the first error information and transmit this information to the second-1 terminal (120). Based on the received user interface display information, the second-1 terminal (120) may display an elapsed time overlay in a numerical form on the time progress bar.
[0096] The user interface displayed on the 2-1 terminal (120) may include emotional emoticons to provide feedback according to the error level. The emotional emoticons may be changed stepwise, such as a smiling face, a blank face, or a frowning face, depending on the error level. Based on the first error information, the server (100) may generate user interface display information that instructs the emotional emoticons to be changed stepwise from 'smiling face → blank face → frowning face' as the error level increases, and transmit this information to the 2-1 terminal (120). The 2-1 terminal (120) may change the display state of the emotional emoticons based on the received user interface display information.
[0097] The server (100) may grant voice input rights to the second-2 terminal (130) or the first terminal (110) based on the first error information.
[0098] The server (100) may grant voice input rights to the second-2 terminal (130) or the first terminal (110) if the number of errors resulting from the second user's reading performance exceeds a preset standard number of errors or exceeds the time limit.
[0099] The 2-2 terminal (130) may be a terminal used by a third user, who is one of the learners participating in the reading group.
[0100] For example, when the server (100) grants voice input rights to the second-2 terminal (130), the second-2 terminal (130) transmits reading data to the server (100) in the same manner as the operation of the second-1 terminal (120) described above, and the server (100) generates error information based on the reading data and, based on this, may again grant voice input rights to the second-1 terminal (120) or the first terminal (110).
[0101] Alternatively, if the server (100) grants voice input rights to the first terminal (110), the first terminal (110) transmits reading data to the server (100) in the same manner as the operation of the second-1 terminal (120) described above, and the server (100) generates error information based on the reading data and, based on this, may again grant voice input rights to the second-1 terminal (120) or the second-2 terminal (130). The first user may intentionally read the e-book text incorrectly to induce participation from the learner. The server (100) may switch voice input rights from the first terminal (110) to the second-1 terminal (120) or the second-2 terminal (130) to induce the learner to participate in reading.
[0102] The server (100) can create one or more virtual users to encourage active participation of learners participating in the reading session and have them participate in the reading session.
[0103] Virtual users are AI-based virtual participants who appear in the participant list in the same way as actual learners and can be given a turn to read according to a set order.
[0104] The server (100) can generate profile information corresponding to a virtual user and provide the profile information to each of the learners' terminals.
[0105] The number of virtual users can be set by the first user when opening a reading meeting, and the probability of error occurrence, error type, difficulty level, etc. can be individually set for each virtual user.
[0106] When it is the virtual user's turn to read aloud, the server (100) can generate a first corrective text, which will be described later, according to the error occurrence probability set for the virtual user, and convert it into a synthesized voice using TTS technology and transmit it to all participants. Through this, actual learners can be provided with a learning opportunity to detect and correct the virtual user's intentional errors.
[0107] To implement an intentional reading error by a virtual user, the server (100) may generate a first modified text that modifies the e-book text provided to the reading meeting.
[0108] The above-mentioned first corrected text may be generated for educational purposes. For example, by reproducing and reading aloud error patterns frequently committed by actual learners, a virtual user can induce actual learners to recognize the errors and learn the correct way to read aloud.
[0109] The above-mentioned first modified text may be generated by the following methods alone or in combination.
[0110] (1) Random change method: Randomly changes words in the e-book text by a preset ratio
[0111] (2) Character insertion method: Inserting arbitrary characters into the e-book text
[0112] (3) Learner error pattern-based approach: Reproduces error patterns frequently committed by actual learners using an accumulated database of learner errors.
[0113] The server (100) can selectively apply the above methods according to the difficulty level and educational goals set for each virtual user.
[0114] For example, if the e-book text provided to the above reading meeting is 'I went to the library with a friend today.' and the preset ratio is '20% of the total number of words', the server (100) says 'I went to the library with a friend today' libraryIn the text 'I went to...', 'library' is randomly changed to 'park' to 'I went to...' with 'I went to...' today with a friend park The above-mentioned first modified text, such as 'I went to', can be generated. (In this example, words are based only on words with substantive meaning, excluding particles.)
[0115] Alternatively, the server (100) may insert additional characters into the e-book text provided to the reading meeting, 'I went to the library with a friend today.', to create 'I went to the library with a friend today' I went The above-mentioned first modified text such as ' can be generated. The method of inserting additional characters into the e-book text provided to the reading meeting can be utilized for the purpose of inducing the learner's concentration.
[0116] Alternatively, the server (100) may generate the first corrected text by modifying the e-book text provided to the reading meeting based on the learner's reading error data. The server (100) may generate the first corrected text by selecting and storing in memory words or sentences that multiple learners frequently make reading errors on, and then modifying the e-book text provided to the reading meeting based on this.
[0117] For example, the server (100) can build an error pattern database by classifying words, sentence structures, and syllable combinations that are difficult to pronounce and that are frequently misread by multiple learners by age group and level, and by selecting error patterns that match the characteristics of the current reading group participants and intentionally reflecting them in the first corrected text, thereby enabling the virtual user to make educationally meaningful errors.
[0118] For example, the server (100) includes an error learning model, and the error learning model can classify and learn words, sentence structures, and syllable combinations that are difficult to pronounce, which are frequently read aloud by multiple learners, by age group and level.
[0119] For example, error learning models can be implemented using CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), Transformer, or a combination thereof. For instance, a CNN-RNN hybrid structure, an LSTM-Attention mechanism, or the latest Transformer-based architecture can be utilized.
[0120] For example, the error learning model may further include a K-means clustering model or a hierarchical clustering model. For example, the error learning model may group learners by age group and level by applying a K-means clustering model or a hierarchical clustering model, and define the error characteristics of each group through a Support Vector Machine (SVM) classification model or a Random Forest ensemble model.
[0121] A first corrected text can be generated based on errors learned by an error learning model. For example, a first corrected text can be generated based on error types learned by an error learning model.
[0122] The server (100) may include an error pattern backtracking model that reverse-generates errors for educational purposes from the correct text. The model may include the following main components.
[0123] (1) Error pattern database: Classifies and stores past learners' reading errors by age and level.
[0124] (2) Error prediction engine: Receives the correct text and learner characteristics as input to predict possible errors
[0125] (3) Error Generator: Selects errors with high educational value from the predicted errors and generates correction text.
[0126] The above educational value can be quantitatively calculated by the following mathematical formula.
[0127] [Mathematical Formula 2]
[0128] Educational Value Score (EV) = w₁ × Frequency Score + w₂ × Learning Transfer Score + w₃ × Difficulty Fit Score
[0129] Here,
[0130] - Frequency Score = (Number of occurrences of the error) / (Total number of error occurrences) × 100
[0131] - Learning Transfer Score = (Similar Error Improvement Rate) × 100
[0132] - Difficulty Appropriateness Score = 1 - |Current Learner Level - Error Difficulty Level| / Maximum Level Difference
[0133] - w₁, w₂, w₃: weights (w₁ + w₂ + w₃ = 1)
[0134] The similar error improvement rate can be a percentage indicator showing how much the incidence rate of other errors of the same category or similar characteristics has decreased after learning a specific type of error. For example, it can measure the transfer learning effect where the incidence rate of other final consonant pronunciation errors (n, d, l, etc.) improves together after learning the final consonant 'ㄱ' pronunciation error.
[0135] For example, the current learner level can be manually set by the teacher (first user) or obtained by a server-based learner level evaluation model. The current learner level is a numerical value that quantitatively evaluates the learner's overall reading ability and can be calculated on a 0 to 10-point scale by synthesizing past learning history, accuracy, fluency, and comprehension. A higher score indicates superior reading ability, and the relative position can be determined by comparing it with the average level for each age group.
[0136] For example, the error difficulty level evaluation model can be manually set by a teacher (first user) or obtained by the server's error difficulty level evaluation model. The error difficulty level may be a value expressing the difficulty of learning and correcting a specific type of error on a scale of 0 to 10. It can be calculated by considering linguistic complexity, cognitive load, and the average acquisition time by age, and a higher score may indicate an error that is more difficult to learn. For example, simple consonant deletion can be classified as low difficulty, while complex grammatical errors can be classified as high difficulty.
[0137] For example, the maximum level difference is the maximum possible difference between the learner level and the error difficulty level, which can mean 10, the theoretical maximum difference when both levels use a 0-10 point scale. This can be used as the denominator for normalization when calculating the difficulty-fit score, ensuring that the result always falls within the 0-1 range.
[0138] In Equation 2, the weights (w₁, w₂, w₃) may be coefficients representing the relative importance of each component (frequency score, learning transfer score, and difficulty suitability score) when calculating the educational value score. The sum of the three weights can always be 1 (w₁ + w₂ + w₃ = 1), and can be adjusted according to educational goals or learner characteristics. For example, w₁ can be set large if frequency is emphasized, and w₂ can be set large if transfer effects are emphasized.
[0139] For example, in the case of lower elementary grades, pronunciation errors of the final consonant 'ㄱ' account for 30% of all errors (occurrence frequency score = 30), and learning this improves similar final consonant pronunciation errors by an average of 40% (learning transfer score = 40), and when the learner level and the difficulty of the error are appropriate (difficulty appropriateness score = 80), if w₁ = 0.3, w₂ = 0.5, and w₃ = 0.2, the educational value score can be calculated as 45 points.
[0140] Errors with an educational value score above the threshold (e.g., 40 points) may be classified as "errors with high educational value" and selected preferentially.
[0141] The above model can operate with the following simple process.
[0142] - Step 1: Identify the age and level of the current learner group
[0143] - Step 2: Select frequent error types from the group
[0144] - Step 3: Apply the selected error type to the correct answer text
[0145] - Step 4: Adjusting the difficulty of generated errors
[0146] For example, since errors in pronunciation of final consonants are frequent among lower elementary students, errors such as changing 'school' to 'school' can occur.
[0147] The educational errors generated through the aforementioned error pattern backtracking model provide the following effects. First, it improves metacognitive abilities by allowing learners to experience in advance the errors they frequently commit. Second, it enhances learning effectiveness through active learning that takes place in the process of listening to and correcting the errors of others. Third, it increases learning duration through gamified error-finding activities.
[0148] The above error pattern backtracking model may be an artificial intelligence model that learns reading error data collected from a large number of learners in the past, predicts errors that actual learners are likely to commit in a specific correct text, and generates a first correct text with high educational value based on this.
[0149] The above error pattern backtracking model can operate based on the following mathematical relationship.
[0150] [Mathematical Formula 3]
[0151] E_pred = f(T_correct, D_historical, C_context)
[0152] Here,
[0153] - E_pred: Set of predicted error patterns {e₁, e₂, ..., e n}
[0154] - T_correct: Correct text (string)
[0155] - D_historical: Past learner error database
[0156] - C_context: Current learning context vector (age, level, learning progress, etc.)
[0157] - f: Deep learning-based backtracking function (RNN or Transformer)
[0158] - e i : i-th prediction error (position, error type, probability)
[0159] Unlike methods utilizing a static error database, the error pattern prediction system based on the above mathematical formula 3 enables the generation of dynamic and contextual errors. By comprehensively considering the correct text (T_correct), historical data (D_historical), and current context (C_context), it is possible to predict educationally meaningful errors rather than simple random errors. Through this, the learning transfer effect can be maximized by providing errors that match the learner's actual weaknesses.
[0160] The above backtracking function f may have a deep learning-based recurrent neural network (RNN) or Transformer structure and can probabilistically predict possible errors at each phoneme, morpheme, and syntactic unit of the input correct text.
[0161] The above-mentioned historical learner error database (D_historical) may be structured as follows. The database may be in the form of a relational database including a learner information table, an error record table, and a text information table. The learner information table stores learner ID, age, gender, learning level, etc., and the error record table records error ID, learner ID, text ID, error type, error location, time of occurrence, etc. The text information table may include text ID, original text content, difficulty level, topic classification, etc.
[0162] The above current learning context vector (C_context) is represented as an n-dimensional real vector, where each dimension is a numerical value of a specific learning characteristic. The above context vector can be composed of a 5-dimensional vector such as C_context = [age, level, progress, difficulty, session_count]. Here, age is a value normalized from the learner's age (0-1), level is a value normalized from the learner's level (0-1), progress is the current curriculum progress rate (0-1), difficulty is the current text difficulty (0-1), and session_count is a value normalized from the cumulative number of learning sessions on a logarithmic scale.
[0163] The deep learning-based backtracking function f described above has a multilayer neural network structure and can be composed of an input layer, a hidden layer, and an output layer. The input layer receives an integrated vector formed by concatenating the embedding vector of the correct text, the feature vector of past error data, and the learning context vector. The hidden layer learns error occurrence patterns at each location within the text through a recurrent structure based on an RNN or Transformer. The output layer can output the probability of occurrence for each error type at each location through a softmax activation function.
[0164] The above i-th prediction error (e i ) can be expressed in the form of a 3-tuple: e i= (position_i, error_type_i, probability_i). position_i is an integer index indicating the location of the error within the text, error_type_i is a category code indicating the type of error (e.g., 1=pronunciation error, 2=grammar error, 3=lexical error, etc.), and probability_i is a real value between 0 and 1 indicating the probability that the error occurs at that location.
[0165] The above error position representation (position_i) can be expressed as a sequential index starting from 0 after dividing the original text into phoneme units or character units. For example, in the case of the text "I am going to school," if divided into character units, it can be indexed as [hak(0), gyo(1), e(2), gap(3), ni(4), da(5)]. Error type classification can be organized hierarchically according to linguistic characteristics and can be subdivided into major categories (pronunciation, grammar, vocabulary), medium categories (consonants, vowels, particles, etc.), and minor categories (specific phonemes, specific endings, etc.).
[0166] The above probability value (probability_i) is a normalized value between 0 and 1 and can be calculated using a sigmoid function or a softmax function. A high probability value (e.g., 0.8 or higher) means that there is a high probability that the error will occur at that location, and a low probability value (e.g., 0.2 or lower) indicates that there is a low probability that the error will occur. The server (100) can determine whether an actual error will occur based on a preset threshold (e.g., 0.5).
[0167] The above error pattern backtracking model may have a multi-stage encoder-decoder structure. The first encoder is a phonetic feature encoder that receives the phoneme sequence of the correct text as input, outputs a pronunciation difficulty vector for each phoneme, and performs the function of analyzing the pronunciation complexity of consonants such as 'ㄱ', 'ㄷ', and 'ㅂ', and vowels such as 'ㅏ', 'ㅓ', and 'ㅗ'. The second encoder is a linguistic feature encoder that receives the results of morphological analysis and syntactic structure as input, outputs a grammatical complexity vector, and analyzes the possibility of errors in particles, ending changes, and spacing. The third encoder is a cognitive load encoder that receives text length, vocabulary difficulty, and sentence complexity as input, outputs a cognitive processing load vector, and analyzes the possibility of errors due to limitations in attention and memory capacity.
[0168] The decoder of the above error pattern backtracking model may include a probabilistic error generation mechanism.
[0169] [Mathematical Formula 4]
[0170] P(error_i | T_correct, position_j) = softmax(W_error · concat(h_phonetic, h_linguistic, h_cognitive) + b_error) i
[0171] Here,
[0172] - P(error_i | T_correct, position_j): Probability of the i-th error type occurring at the j-th position
[0173] - concat(): Vector concatenation operation
[0174] - h_phonetic ∈ : Phonetic characteristic vector
[0175] - h_linguistic ∈ : Linguistic characteristic vector
[0176] - h_cognitive ∈ : Cognitive load vector
[0177] - W_error ∈ : Error prediction weight matrix (n: number of error types)
[0178] - b_error ∈ : Bias vector
[0179] - softmax() i : i-th element of the softmax function
[0180] The probabilistic error generation mechanism of the above Equation 4 implements the naturalness of error occurrence through multi-layered characteristic analysis. Through a softmax probability distribution that comprehensively considers phonetic, linguistic, and cognitive characteristics, errors can be generated in the order of those most likely to be committed by actual learners. This probability-based approach is effective in preventing predictable error patterns and maintaining the learner's continuous concentration and vigilance.
[0181] The above vector concat operation may be a linear transformation operation that combines multiple vectors into a single unified vector. For example, three d-dimensional vectors h_phonetic ∈ , h_linguistic ∈ , h_cognitive ∈ Given, concat(h_phonetic, h_linguistic, h_cognitive) is a 3d-dimensional consolidation vector [h_phonetic; h_linguistic; h_cognitive] ∈ It can generate. In this case, the semicolon (;) can mean connecting the vectors vertically.
[0182] The above phonetic feature vector (h_phonetic) may be a d-dimensional real vector that quantifies the phonetic characteristics of the text. The vector may include consonant complexity, vowel complexity, syllable structure complexity, the presence or absence of final consonants, the possibility of assimilation, etc. For example, when d=64, h_phonetic = [c₁, c₂, ..., c64 It can be expressed as ], and each element c i The intensity of a specific phonetic characteristic can be represented by a value between 0 and 1. Consonant complexity can be set to a low value for 'ㄱ, ㄴ' and a high value for 'ㅊ, ㅋ' depending on the difficulty of pronunciation.
[0183] The above linguistic feature vector (h_linguistic) may be a d-dimensional real vector that quantifies the grammatical and semantic features of the text. The vector may include part-of-speech information, particle complexity, ending complexity, syntactic structure complexity, lexical frequency, etc. For example, in the case of the phrase 'school-e', it is analyzed as a combination of the noun 'school' and the particle 'e', so the part-of-speech information can be one-hot encoded as [1, 0, 0, 1, 0, ...], and the particle complexity can be set to a value such as 0.3 depending on the usage frequency and grammatical complexity of 'e'.
[0184] The above cognitive load vector (h_cognitive) may be a d-dimensional real vector that quantifies the amount of cognitive resources required for text processing. The vector may include word length, sentence length, vocabulary frequency, semantic abstraction, working memory load, etc. For example, since longer words require more working memory, the value of the corresponding dimension may increase linearly as word length increases. Semantic abstraction may be set to a low value for concrete nouns (e.g., 'apple') and a high value for abstract concepts (e.g., 'definition').
[0185] The error prediction weight matrix (W_error) may be a learnable parameter matrix that converts integrated feature vectors into error prediction scores. The size of the matrix may be n × 3d, where n is the total number of error types and 3d is the dimension of the connected feature vectors. The matrix may be learned through a backpropagation algorithm, and initial values may be set using Xavier initialization or He initialization methods. Each row may represent a weight for a specific error type, and through learning, greater weights may be assigned to features that are highly relevant to that error type.
[0186] The number of error types (n) mentioned above may represent the total number of error categories predictable by the system. The error types may follow a hierarchical classification system and may consist of 6 major categories (pronunciation errors, intonation errors, grammar errors, vocabulary errors, fluency errors, comprehension errors), 24 medium categories, and 72 minor categories, resulting in a total of n=102 error types. Each error type may be identified by a unique integer index (from 0 to n-1) and may be represented using one-hot encoding.
[0187] The above bias vector (b_error) may be a learnable bias parameter vector added to the linear transformation of the neural network. The magnitude of the vector may be n-dimensional, and each element may represent the underlying tendency of the corresponding error type to occur. The bias vector provides an underlying probability offset for each error type regardless of the input, thereby allowing the model to reflect the tendency for specific error types to occur more frequently overall. The initial value may generally be set to 0 and may be adjusted based on the data during the training process.
[0188] The above decoder outputs the probability of each error type that can occur at each text location, thereby enabling the selective generation of errors with the highest educational value. An algorithm for selectively generating errors with the highest educational value (error selection algorithm) may be as follows.
[0189] The error selection algorithm can be implemented through the following five-step process.
[0190] Step 1 (Generation of Candidate Errors): The server inputs the correct text and the learner's age group information into the error pattern database to search for a list of candidate errors that may occur in that age group. For example, when a second-grade elementary student reads the word "school," possible errors that may occur include "hakgo" (dropping of the final consonant), "hagyo" (initial consonant error), and "hakgyo-eul" (adding of the particle), which can be extracted as candidates.
[0191] Step 2 (Calculation of Educational Value Score): The server calculates the frequency score, learning transfer score, and difficulty fit score for each candidate error, and then calculates the educational value score by applying the above mathematical formula 4. The frequency score reflects the statistical frequency of a specific error occurring in the corresponding age group, the learning transfer score indicates the degree of improvement in other similar errors when the error is learned, and the difficulty fit score is a value that evaluates the suitability of the error difficulty to the current learner's level.
[0192] Step 3 (Threshold-based filtering): The server selects only errors whose calculated educational value score is above a preset threshold (e.g., 40 points). This is a process designed to exclude errors with minimal educational effect and select only those errors that are of practical help to learners.
[0193] Step 4 (Ensuring Diversity by Error Type): The server classifies the selected errors by type, such as pronunciation errors, grammar errors, and vocabulary errors, and then selects the errors with the highest scores so as not to exceed the maximum number of selections (e.g., 2) for each type. This is intended to prevent the concentrated occurrence of specific types of errors and to allow learners to experience a variety of error types.
[0194] Step 5 (Final Error Selection): If the total number of selected errors exceeds the maximum allowed number of errors, the server selects the final errors in order of highest educational value score. Conversely, if the number of selected errors is less than or equal to the maximum allowed number, all selected errors are used.
[0195] For example, when the algorithm is applied to a second-grade elementary student on the correct answer text "I go to school," candidates such as "I go to school" (final consonant omission), "I go to school" (particle error), and "I go home" (initial consonant error) are generated in the first stage, and in the second stage, educational value scores of 52, 48, and 35 are calculated respectively, and in the third stage, the first two errors exceeding a threshold of 40 points are selected, so that one of "I go to school" or "I go to school" can be finally generated as the first corrected text.
[0196] For example, the step of generating the first corrected text may include the step of searching for candidate errors corresponding to the age group of the learner from an error pattern database; the step of calculating an educational value score for each candidate error, including a frequency score, a learning transfer score, and a difficulty suitability score; the step of selecting errors whose educational value score is greater than or equal to a preset threshold; and the step of generating the first corrected text by applying the selected errors to the e-book text.
[0197] The error data used to train the aforementioned error pattern backtracking model may feature a multidimensional labeling system. In terms of error type, errors are classified into phonetic errors (consonant deletion, consonant addition, vowel change, syllable omission, etc.), linguistic errors (particle errors, ending errors, spacing errors, word order errors, etc.), and cognitive errors (repetition errors, ellipsis errors, substitution errors, insertion errors, etc.). In terms of difficulty, errors are categorized into Level 1 (clear errors immediately recognizable), Level 2 (subtle errors recognizable through careful listening), and Level 3 (advanced errors requiring specialized knowledge). In terms of educational effectiveness, errors are classified into Effect A (errors with immediate learning transferability), Effect B (errors improved through repetitive learning), and Effect C (errors contributing to long-term language proficiency improvement).
[0198] The above multidimensional labeling system consists of three independent dimensions, and each error data can be represented as a 3-dimensional coordinate value. For example, an error of pronouncing "school" as "hakgo" can be labeled as (phonetic error-consonant deletion, level 1, effect A) and represented as a coordinate value (1-1, 1, A).
[0199] The above error type dimension can be subdivided as follows.
[0200] - Phonetic Errors (Code 1): Consonant deletion (1-1), Consonant addition (1-2), Vowel change (1-3), Syllable omission (1-4)
[0201] - Linguistic Errors (Code 2): Particle Errors (2-1), Ending Errors (2-2), Spacing Errors (2-3), Word Order Errors (2-4)
[0202] - Cognitive Errors (Code 3): Repetition Error (3-1), Omission Error (3-2), Substitution Error (3-3), Insertion Error (3-4)
[0203] The above difficulty dimensions can be classified as follows based on the clarity of error recognition.
[0204] - Level 1: Immediately recognizable errors within 3 seconds of detection
[0205] - Level 2: Errors requiring careful listening with a detection time of 3 to 10 seconds
[0206] - Level 3: Errors requiring professional analysis with a detection time exceeding 10 seconds
[0207] The above dimensions of educational effectiveness can be classified as follows based on the continuity of learning transfer.
[0208] -Effect A: Immediate transfer error improved by over 90% with a single training session
[0209] - Effect B: Gradual transfer error improved with 3 to 5 learning iterations
[0210] -Effect C: Cumulative transfer error requiring long-term learning of 10 or more iterations
[0211] The server (100) can generate customized errors for each learner by utilizing the above multidimensional labeling system. For example, it can first provide errors corresponding to the coordinates (1-1, Level 1, Effect A) to beginner learners and provide errors corresponding to the coordinates (2-3, Level 3, Effect C) to advanced learners.
[0212] For example, the step of generating the first correction text may include the step of classifying error data stored in an error pattern database according to a multidimensional labeling scheme.
[0213] For example, the above multidimensional labeling system may include a first dimension (error classification dimension) classified into phonetic errors, linguistic errors, and cognitive errors, a second dimension (difficulty dimension) classified into levels 1 to 3 according to error detection time, and a third dimension (educational effect dimension) classified into effects A to C according to learning transfer characteristics.
[0214] For example, the step of generating the first corrected text may further include the step of selectively extracting errors corresponding to a specific coordinate range of the multidimensional labeling system based on the age group and level of the learners participating in the reading session.
[0215] By applying the aforementioned multidimensional labeling system, the following effects can be achieved. First, error data expressed as three-dimensional coordinate values allows for the precise analysis of individual learner weaknesses, enabling the design of personalized learning paths. Second, error selection considering both difficulty and educational effectiveness is possible, ensuring the most effective learning at the learner's current level. Third, the accumulated three-dimensional data enables statistical analysis of error patterns by age and level, which can be utilized for establishing educational policies.
[0216] For example, the above specific coordinate range may be set to prioritize the selection of errors where the difficulty dimension is Level 1 and the educational effect dimension is Effect A (for beginner learners), errors where the difficulty dimension is Level 2 and the educational effect dimension is Effect B (for intermediate learners), and errors where the difficulty dimension is Level 3 and the educational effect dimension is Effect C (for advanced learners).
[0217] The server (100) can apply data augmentation techniques to improve the learning performance of the error pattern backtracking model. In phonetic augmentation, errors based on standard language are modified to match regional pronunciation characteristics through regional dialect modification, and error patterns are generated that consider differences in articulation ability by age group by reflecting pronunciation characteristics by age. In linguistic augmentation, errors occurring during the conversion process between spoken and written language are simulated, and modified text is generated that adjusts sentence length and structural complexity. In contextual augmentation, error patterns of specialized vocabulary by topic, such as science, history, and literature, are reflected, and error characteristics of beginner, intermediate, and advanced learners at each stage are considered.
[0218] The specific operation process of the above error pattern backtracking model is as follows. In Step 1, the input text is decomposed into phonemes, and the pronunciation complexity for each phoneme is calculated. For example, "학교에 갑습니다" can be decomposed into [ㅎ,ㅏ,ㄱ,ㄱ,ㅛ,ㅇ,ㅔ,ㄱ,ㅏ,ㅂ,ㄴ,ㅣ,ㄷ,ㅏ], and the complexity of the final consonant 'ㄱ' can be calculated as 0.8, and the complexity of the diphthong 'ㅛ' as 0.6. In Step 2, error candidates are generated, such as phonetic candidates ("학고에 갑습니다," "하교에 갑습니다"), linguistic candidates ("학교애 갑습니다," "학교에가 갑습니다"), and cognitive candidates ("학교에... 갑습니다," "학교에 갑니까"). In Step 3, an educational effectiveness index is calculated for each error candidate, and the goodness of fit is evaluated considering the characteristics of the current learner group. In Step 4, the error with the maximum educational effectiveness index is finally selected.
[0219] The above error pattern backtracking model may include a real-time adaptation mechanism. In the feedback collection phase, learners' error detection times are measured, the number of correction attempts and success rates are recorded, and changes in engagement are monitored. In the model update phase, real-time weight adjustments are performed through an online learning algorithm, model corrections are made for learner responses that differ from expectations, and training data is automatically added when a new error pattern is discovered. As an adaptation strategy, if the error detection time is less than the lower threshold, the next error difficulty is increased by 0.1 to generate a more difficult error, and if the error detection time is greater than the upper threshold, the next error difficulty is decreased by 0.1 to generate an easier error.
[0220] The performance of the aforementioned error pattern backtracking model can be evaluated using the following metrics. Accuracy metrics include error prediction accuracy, which is the concordance rate between actual learner errors and predicted errors, and difficulty fit, which is the suitability of the difficulty of the generated errors to the learner level. Educational effectiveness metrics include learning improvement, which is the degree of improvement in the accuracy rate before and after error learning; transfer effect, which is the improvement in generalization ability to similar errors; and motivation, which is the change in learner engagement and concentration. Efficiency metrics include processing speed required for real-time error generation, memory usage required for model operation, and scalability, which is the change in performance due to an increase in the number of concurrent users.
[0221] For example, the error pattern backtracking model may include a multi-stage encoder that decomposes the correct text into phonetic characteristics, linguistic characteristics, and cognitive load characteristics; a probabilistic decoder that calculates the probability of an error occurring at each text position based on the output of the multi-stage encoder; and an error selection module that selects an optimal error pattern by combining the probability of an error occurring and a predefined educational effect index.
[0222] For example, the above error pattern backtracking model may further include an adaptive learning module that collects real-time response data from learners and dynamically updates model parameters through an online learning algorithm.
[0223] The server (100) can generate first synthetic voice data by performing text-to-speech (TTS) on the first modified text and transmit it to the first terminal (110) and the second-1 terminal (120).
[0224] Synthesized speech data refers to audio data generated by the server (100) converting the content of text into speech through text-to-speech conversion.
[0225] The server (100) may be equipped with deep learning-based text-to-speech conversion models such as Tacotron2, FastSpeech2, and VITS.
[0226] The server (100) can generate second error information based on the first synthesized voice data and grant voice input rights to the first terminal (110) or the second-1 terminal (120) based on the second error information.
[0227] The server (100) can analyze the first synthetic voice data generated based on the first modified text to detect reading errors intentionally included in the first synthetic voice. When the server (100) detects an error by a virtual user, the server (100) can grant voice input rights to actual learner terminals (e.g., terminal 2-1, terminal 2-2, etc.) so that actual learners can point out the error or demonstrate the correct reading.
[0228] The server (100) may grant voice input rights to the first terminal (110) or the second-1 terminal (120) based on the second error information. Learners participating in the reading session may be repeatedly granted opportunities to read aloud while maintaining their interest in learning to read aloud.
[0229] The server (100) can automatically adjust the division unit of the e-book text according to the pre-set class time (by the first user, etc.). For example, if the class is set to 30 minutes, the number of divisions per sentence unit can be set relatively small, and if the class is set to 60 minutes, the number of divisions per sentence unit can be set relatively large.
[0230] The server (100) monitors the elapsed time in real time during the class and can send a notification of the end to the participants of the reading group when the set class time approaches. For example, at 90% of the set class time, a message such as "The class will end soon" can be sent to each terminal.
[0231] The server (100) can automatically end the reading session when the set class time has elapsed and generate summary information of the learning results of the session and provide it to each participant's terminal. The summary information of the learning results may include statistical data such as the individual participant's reading participation time, number of errors, and accuracy.
[0232] The server (100) can provide customized content based on the class time setting. For example, for a 30-minute class, it can provide text consisting mainly of short sentences that are easy to maintain concentration, and for a long class of 60 minutes or more, it can provide text of varying difficulty levels in stages to maximize the learning effect.
[0233] Figure 2 is a conceptual diagram showing the server of Figure 1.
[0234] As illustrated in FIG. 2, the server (100) can perform calculations based on user input or based on a predetermined algorithm as a computing device.
[0235] For example, the server (100) may include a processor (101), memory (102), and a communication module (103).
[0236] The processor (101) can control other configurations by executing instructions stored in memory (102). The processor (101) can execute instructions stored in memory (102).
[0237] A processor (101) is a component capable of performing calculations and controlling other devices. It may primarily refer to a central processing unit (CPU), an application processor (AP), a graphics processing unit (GPU), etc. Additionally, a CPU, AP, or GPU may include one or more cores within it, and the CPU, AP, or GPU may operate using operating voltage and clock signals. However, while a CPU or AP may consist of a few cores optimized for serial processing, a GPU may consist of thousands of smaller and more efficient cores designed for parallel processing.
[0238] The processor (101) can provide or process appropriate information or functions to the user by processing signals, data, information, etc. that are input or output through the components described above, or by running an application stored in memory (102).
[0239] The memory (102) stores data that supports various functions of the server (100). The memory (102) can store a number of applications (application programs or applications) running on the computing device, data for the operation of the computing device, and instructions. At least some of these applications may be downloaded from an external server via wireless communication. Additionally, the applications may be stored in the memory (102), installed on the computing device, and driven by the processor (101) to perform the operation (or function) of the computing device.
[0240] The memory (102) may include at least one type of storage medium among flash memory type, hard disk type, SSD type (Solid State Disk type), SSD type (Silicon Disk Drive type), multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), RAM (random access memory; RAM), SRAM (static random access memory), ROM (read-only memory; ROM), EEPROM (electrically erasable programmable read-only memory), PROM (programmable read-only memory), magnetic memory, magnetic disk, and optical disk. Additionally, the memory (102) may include web storage that performs storage functions over the internet.
[0241] The communication module (103) performs the transmission and reception of information with a base station or other components including communication functions through an antenna. The communication module (103) may include a modulation unit, a demodulation unit, a signal processing unit, etc. For example, the communication module (103) may perform wireless communication functions and / or wired communication functions.
[0242] Wireless communication can refer to communication using communication facilities installed by telecommunication companies and a wireless communication network that uses the frequency of said communication facilities. In this case, the communication module (103) can be used in various wireless communication systems such as CDMA (code division multiple access), FDMA (frequency division multiple access), TDMA (time division multiple access), OFDMA (orthogonal frequency division multiple access), SC-FDMA (single carrier frequency division multiple access), etc. Furthermore, the communication module (103) can also be used in 3GPP (3rd generation partnership project) LTE (long term evolution), etc. In addition, not only 5G communication currently being commercialized but also 6G, which is scheduled for future commercialization, can be used. However, the present specification can utilize an existing communication network without being restricted by such wireless communication methods.
[0243] In addition, short-range communication technologies such as Bluetooth, BLE (Bluetooth Low Energy), Beacon, RFID (Radio Frequency Identification), NFC (Near Field Communication), Infrared Data Association (IrDA), UWB (Ultra Wideband), and ZigBee can be used.
[0244] Based on the above-described contents, a relay reading education method according to a preferred embodiment of the present specification will be described in detail as follows.
[0245] FIG. 3 is a flowchart schematically illustrating a relay reading education method according to an embodiment of the present invention, FIG. 4 is a flowchart illustrating in more detail the step (S120) of providing e-book text of FIG. 3, FIG. 5 is a flowchart illustrating in more detail the step (S140) of generating first error information of FIG. 3, and FIG. 6 is a flowchart illustrating in more detail the step (S150) of transmitting first time information of FIG. 3.
[0246] For reference, the entity performing the relay reading education method of Fig. 3 may be the aforementioned server or the processor of the server.
[0247] Referring to FIG. 3, the relay reading education method may include the steps of: receiving a request to open a reading session from a first terminal (S110); providing an e-book text to the reading session (S120); granting voice input rights to the second-1 terminal (S130); receiving first reading data from the second-1 terminal and generating first error information based on the first reading data (S140); transmitting first time information to the second-1 terminal based on the first error information (S150); and granting voice input rights to the second-2 terminal or the first terminal based on the first error information (S160).
[0248] Referring to FIG. 4, the step (S120) may include receiving a request to provide e-book text from the first terminal (S1210), providing a pre-stored list of e-books to the first terminal (S1220), receiving information about an e-book selected from the list of e-books from the first terminal (S1230), and dividing the entire text constituting the selected e-book into sentence units and then sequentially providing a pre-set number of sentences to the reading group (S1240).
[0249] Referring to FIG. 5, the step (S140) may include receiving the first reading data from the second-1 terminal (S1410), applying the first reading data to a pre-trained speech recognition model to convert it into a first speech recognition text (S1420), and generating the first error information based on the first speech recognition text and the e-book text provided to the reading group (S1430).
[0250] The above first error information may include the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit.
[0251] Referring to FIG. 6, step (S150) may include a step (S1510) of generating user interface display information based on the first error information, a step (S1520) of transmitting the user interface display information to the second-1 terminal, and a step (S1530) of controlling the state of at least one of the graphic elements constituting the user interface displayed on the second-1 terminal to change according to the user interface display information. At this time, the user interface display information may include at least one of i) information instructing that some of the pre-displayed multiple segments disappear as the number of errors increases, ii) information instructing that the display state of the pre-displayed time progress bar be changed according to the elapsed time, and iii) information instructing that the display form of the pre-displayed icon be changed according to the error level (the error level means a level calculated based on the number of errors and the elapsed time).
[0252] Figure 7 is a flowchart schematically illustrating a relay reading education method that includes the participation of artificial intelligence in reading.
[0253] Referring to FIG. 7, the relay reading education method comprises the steps of: receiving a request to open a reading session from a first terminal (S110); providing an e-book text to the reading session (S120); granting voice input authority to the second-1 terminal (S130); receiving first reading data from the second-1 terminal and generating first error information based on the first reading data (S140); transmitting first time information to the second-1 terminal based on the first error information (S150); granting voice input authority to the second-2 terminal or the first terminal based on the first error information (S160); generating a first modified text by modifying the e-book text provided to the reading session (S170); generating first synthesized voice data by performing text-to-speech (TTS) on the first modified text (S180); and transmitting the first synthesized voice data to the first terminal and the second-1 It may include a step of transmitting to a terminal (S190), a step of generating second error information based on the first synthesized voice data (S200), and a step of granting voice input authority to the first terminal or the second-1 terminal based on the second error information (S210).
[0254] The first modified text above refers to a text generated by at least one of the following text generation methods: a method of randomly changing the e-book text provided to the reading meeting by a preset ratio, or a method of additionally inserting arbitrary characters into the e-book text provided to the reading meeting.
[0255] For reference, any content in the descriptions of FIGS. 3 to 7 that is identical or duplicates the content described in FIGS. 1 to 2 may be omitted. A person skilled in the art can easily understand the technical concept of FIGS. 3 to 7 based only on the descriptions of FIGS. 1 to 2.
[0256] For example, the server (100) may provide an interface that allows each setting value of the scoring system to be individually adjusted through the first terminal (110). The first user may adjust the basic score increase amount, score increase time interval, fever time activation time, fever time score increase amount, and allowed number of errors according to the characteristics of the reading group.
[0257] For example, the first user can set the number of allowed errors to be increased to 3 for the beginner learner group and the fever time activation time to be shortened to 90 seconds so that learners can easily feel a sense of accomplishment. Conversely, for the advanced learner group, the number of allowed errors can be reduced to 1 and the basic score increase amount lowered to 5 points to create a more challenging environment.
[0258] For example, the server (100) can adjust the score increase time interval to 5 seconds, 10 seconds, 15 seconds, etc., and can set the basic score increase amount to 5 points, 10 points, 15 points, 20 points, etc. The fever time activation time can be adjusted to 60 seconds, 90 seconds, 120 seconds, 180 seconds, etc., and the fever time score increase amount can be set to 1.5 times, 2 times, 2.5 times, 3 times the basic score increase amount.
[0259] For example, the server (100) can calculate each participant's score in real time and display it on all participant terminals. The score information may include the current score, cumulative score, fever time activation status, remaining allowed error count, etc.
[0260] For example, the server (100) can calculate the final score ranking at the end of the reading session and provide reward information (e.g., virtual badge, title, etc.) for the excellent participants. This can stimulate the learners' competitive spirit and motivation to achieve.
[0261] For example, the server (100) may provide a mission assignment and reward system to induce review learning by the learner at the end of the reading session and to promote continuous learning participation.
[0262] For example, when a reading session ends, the server (100) may receive review mission setting information from the first terminal (110). The review mission setting information may include information on a specific page range of the e-book studied in the reading session and information on the number of reviews.
[0263] For example, the first user can input the number of pages that need to be reviewed among the e-book texts learned in the reading group through the first terminal (110). For example, the first user can input a page range such as "from page 15 to page 20".
[0264] For example, the server (100) can generate an automatic notification message based on the review mission setting information and send it to a parent terminal associated with each learner. The automatic notification message can be sent in the form of SMS, KakaoTalk notification, email, or a dedicated application push notification.
[0265] For example, the above automatic notification message may include specific review instructions in the form of, "Please read pages 15 through 20 of [e-book title] learned in today's reading class twice."
[0266] For example, the server (100) can basically set the number of reviews to 2 times, but can be adjusted to 1 time, 2 times, 3 times, etc. according to the settings of the first user. In addition, the deadline for completing the review can be set to before the start of the next reading meeting.
[0267] For example, the server (100) may operate an authentication system to verify whether the learner has completed the mission. The learner may perform review completion authentication through the 2-1 terminal (120) or the 2-2 terminal (130), and this may be done in the form of voice recording, taking a photo, or answering a simple quiz.
[0268] For example, the server (100) may provide a visual reward to a learner who has completed a mission. The visual reward may be a decorative element applied to the learner's profile character, and may be in the form of a laurel wreath, a crown, a badge, a star decoration, etc.
[0269] For example, the server (100) can control the application of the visual reward to the user interface displayed on the terminal of the learner who completed the mission when the next reading session begins. For example, a laurel wreath may be displayed around the profile image of the learner who completed the mission, or a special icon may be displayed next to the name of the learner in the participant list.
[0270] For example, the server (100) may provide additional rewards to learners who complete missions consecutively. For example, special decorative elements such as a gold laurel wreath for completing missions three times in a row, and a diamond crown for completing missions five times in a row may be provided.
[0271] The above mission and reward system provides the following long-term learning effects. First, review missions prevent memory loss according to the Ebbinghaus forgetting curve, thereby improving the conversion rate of learning content into long-term memory. Second, visual rewards (laurel wreaths, crowns, etc.) satisfy the need for social recognition, increasing the rate of voluntary participation in learning. Third, cumulative rewards for completing consecutive missions help form study habits, and if continued for a certain period, have the effect of solidifying the habit of close reading.
[0272] For example, the server (100) may generate mission completion statistics and provide them to the first terminal (110). The mission completion statistics may include mission completion rates for each learner, consecutive completion records, and the average completion rate of all participants, so that the first user can monitor the review learning status of the learners.
[0273] Figures 8 and 9 are exemplary drawings showing a user interface displayed during a relay reading meeting.
[0274] FIG. 8 is an example of a user interface that visually displays the state in which a reading teacher is currently performing a reading and student 2 is waiting as the next reading runner.
[0275] Figure 9 is an example of a user interface that visually displays the state in which Student 2 is currently performing a reading and the artificial intelligence is set as the next reader.
[0277] Method for classifying types of reading errors
[0279] Hereinafter, based on the above-described contents, a method for classifying reading error types according to a preferred embodiment of the present specification will be described in detail as follows.
[0280] For reference, the entity performing the reading error type classification method, which is an embodiment of the present invention, may be the server (100) described above or the processor (101) of the server (100).
[0281] FIG. 10 is a flowchart schematically illustrating a method for classifying reading error types according to an embodiment of the present invention.
[0282] For reference, the user terminal described below may refer to one of the first terminal (110), the second-1 terminal (120), or the second-2 terminal (130), and the description of the user terminal may be replaced by the description of the first terminal (110), the second-1 terminal (120), or the second-2 terminal (130) described above.
[0283] As illustrated in FIG. 10, the reading error type classification method may include the step (S310) of providing a reference text to a user terminal.
[0284] The above reference text may be text in sentence units extracted from news articles, speeches, everyday conversations, etc., and may refer to the aforementioned e-book text.
[0285] The method for classifying reading error types may include the step (S320) of receiving reading voice data from the user terminal.
[0286] The above reading voice data may be in audio file formats such as WAV, MP3, FLAC, etc.
[0287] The above-mentioned reading voice data may include everyday noise (e.g., keyboard typing sounds, home appliance operation sounds, etc.) or background noise (e.g., car horn sounds, white noise, etc.).
[0288] The method for classifying reading error types may include a step (S330) of preprocessing the reading voice data.
[0289] Step (S330) may include a step of performing noise reduction filtering to pass a frequency band pre-set as where human voice is distributed and attenuate the remaining frequency components.
[0290] The preset frequency band may be set to a range of approximately 80 Hz to 4 kHz, where the main frequencies of human voice are distributed.
[0291] Noise reduction filtering can be performed using band-pass filters, Kalman filters, adaptive filters, etc.
[0292] The method for classifying reading error types may include the step (S340) of generating voice recognition text by applying the preprocessed reading voice data to a preset voice recognition model.
[0293] Speech recognition text may refer to string data in which certain speech data is converted into the form of text by the speech recognition model.
[0294] The above speech recognition model may be a STT (Speech-to-Text) model having a deep learning structure based on CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), and Transformer.
[0295] The method for classifying reading error types may include the step (S350) of generating reading error characteristic information based on the preprocessed reading voice data and the voice recognition text.
[0296] Reading error characteristic information may refer to information used to determine reading errors, such as the accuracy, fluency, speed, and silent intervals of the spoken voice.
[0297] Step (S350) may include at least two of the following operations: a first operation of calculating an edit distance between the reference text and the voice recognition text and setting it as a first characteristic value; a second operation of analyzing a change in pitch in the recited voice data and setting it as a second characteristic value; a third operation of measuring the maximum duration of a continuous silent interval included in the recited voice data and setting it as a third characteristic value; a fourth operation of setting the recited voice data to '1' if the recited voice data is not received within a preset waiting time from the time the reference text is provided to the user terminal, and setting the fourth characteristic value to '0' if the recited voice data is received within the waiting time; a fifth operation of measuring the recitation speed from the preprocessed recited voice data and setting it as a fifth characteristic value; and a sixth operation of dividing the total time recognized as silent in the preprocessed recited voice data by the total playback time and setting it as a sixth characteristic value, and defining the result value from the performed operations as the recitation error characteristic information.
[0298] Edit distance refers to the minimum number of edit operations (insertion, deletion, replacement) required to make two strings identical, and string similarity measurement methods such as the Levenstein distance algorithm can be utilized.
[0299] For example, if the reference text above is 'I go to school' and the speech recognition text above is 'I go to school', since a pronunciation error occurred from 'school' to 'school', the edit distance is calculated as '1' by the above first operation, and the corresponding value can be set as the above first characteristic value.
[0300] For example, if the range of pitch variation in the preprocessed reading voice data is analyzed and measured as '±40Hz', then '40Hz' can be set as the second characteristic value by the second operation.
[0301] For example, if, as a result of analyzing the sections without speech in the preprocessed reading voice data, there exists a continuous silent section of up to '1.5' seconds, '1.5 seconds' may be set as the third characteristic value by the third operation.
[0302] For example, if the above waiting time is '1 minute' and the above reading voice data is not received at all within '1 minute' after the above reference text is provided to the user terminal, the above fourth characteristic value may be set to '1' by the above fourth operation.
[0303] For example, if the above waiting time is '1 minute' and the above reading voice data is received after '40 seconds' after the above reference text is provided to the above user terminal, the above fourth characteristic value may be set to '0' by the above fourth operation.
[0304] Reading speed can refer to the number of characters spoken per unit of time (second).
[0305] For example, if the analysis of the above-mentioned preprocessed voice data reveals that a total of 90 characters were spoken over 10 seconds, '9' may be set as the fifth characteristic value by the above-mentioned fifth operation.
[0306] For example, if the total time recognized as silent in the above-mentioned preprocessed reading voice data is '5 seconds' and the total playback time is '20 seconds', the value calculated by dividing the total time recognized as silent in the above-mentioned preprocessed reading voice data by the total playback time by the above-mentioned 6th operation (5 ÷ 20 = 0.25) can be set as the above-mentioned 6th characteristic value.
[0307] For example, if, as a result of performing all of the above first to sixth operations, the first characteristic value is set to '1', the second characteristic value to '40Hz', the third value to '1.5 seconds', the fourth characteristic value to '0', the fifth characteristic value to '9', and the sixth characteristic value to '0.25', the present invention may generate a feature vector [1, 40, 1.5, 0, 9, 0.25] having the first to sixth characteristic values as components, and define the vector as the reading error characteristic information.
[0308] The method for classifying reading error types may include the step (S350) of inputting the reading error characteristic information into a pre-trained error classification model and outputting at least one of a plurality of preset reading error types.
[0309] The above multiple types of reading errors may include pronunciation errors, intonation errors, reading pause errors, non-response errors, reading speed errors, or concentration loss errors.
[0310] The above error classification model may be an artificial intelligence-based classification model trained to receive the reading error characteristic information, calculate a probability for each of the plurality of reading error types, and select and output all reading error types for which the calculated probability is greater than or equal to a preset threshold.
[0311] For example, the above error classification model is an artificial intelligence model trained using a training dataset that includes the reading error characteristic information and corresponding reading error type labels, and can be implemented as a multi-layer perceptron (MLP)-based classification model.
[0312] The above error classification model can convert input reading error characteristic information into probability values between 0 and 1 by applying a sigmoid function for each reading error type. The above error classification model can produce different probability values by using the same sigmoid function for each reading error type but applying different weights and biases.
[0313] For example, if the preset threshold is '0.7' and the reading error characteristic information is [1, 40, 1.5, 0, 9, 0.25], the error classification model receives the reading error characteristic information and applies a sigmoid function to each reading error type to calculate probability values such as (pronunciation error: 0.82, intonation error: 0.25, reading interruption error: 0.12, non-response error: 0.10, reading speed error: 0.74, concentration loss error: 0.20), and then outputs the 'pronunciation error' (0.82) and 'reading speed error' (0.74) which are greater than or equal to the threshold (0.7) as the final reading error types.
[0314] For example, pre-established thresholds may have differentiated values based on age group. For instance, different thresholds may be set for each type of error. Alternatively, thresholds may be set differently for each age group and error type. As a result, adaptive error assessment tailored to the learner's developmental stage can be performed.
[0315] For example, the age range of learners may be distinguished into a first age range to an n-th age range. For example, the first age range may be the age of first and second graders in elementary school, the second age range may be the age of third and fourth graders in elementary school, and the third age range may be the age of fifth and sixth graders in elementary school. However, such age range distinctions are merely examples, and the scope of rights of this specification is not limited thereto.
[0316] The threshold can be set according to age characteristics. For example, looking at the pronunciation error judgment process, the threshold for the first age group may be 0.8, the threshold for the second age group may be 0.5, and the threshold for the third age group may be 0.3.
[0317] The threshold value can be set according to the characteristics of the error type. For example, since it may be more difficult to pronounce accurately in the first age group, the threshold value for judging pronunciation errors in the first age group may be 0.8, and the threshold value for judging intonation errors or low concentration errors may be 0.7.
[0318] The threshold values may be set according to age characteristics and error type characteristics. For example, for the first age group, the threshold for judging pronunciation errors may be 0.8, and the threshold for judging intonation errors or low concentration errors may be 0.7. In this case, for the third age group (since reading speed may be more important), the threshold for judging pronunciation errors, non-response errors, and low concentration errors for the third age group may be 0.4, and the threshold for judging reading speed errors may be 0.3.
[0319] The threshold value may be automatically determined by the server (100) by a pre-entered algorithm. Alternatively, the threshold value may be manually adjusted by the first user. When the first user opens a reading session, an interface for adjusting the threshold value may be provided to the first terminal to adjust the threshold value.
[0320] By subdividing the threshold values, the learner's reading performance ability can be evaluated.
[0321] For example, in the case of pronunciation errors, the threshold can be subdivided into multiple ranges as follows.
[0322] (1) Level 1: Standard range 0 ~ 0.2
[0323] (2) Level 2: Standard range 0.2 ~ 0.4
[0324] (3)3 Level: Standard range 0.4 ~ 0.6
[0325] (4) Level 4: Standard range 0.6 ~ 0.8
[0326] (5) Level 5: Standard range 0.8 ~ 1.0
[0327] The five levels described above are examples, and more detailed levels may be set depending on the case. For example, the above standard range may be divided from level 1 to level n.
[0328] The first characteristic value, which is the output value of the sigmoid function, may be included in any one of the ranges from level 1 to level n (e.g., level 1 to level 5). For example, the first user may adjust the threshold for pronunciation errors during the next reading based on the range to which the first characteristic value belongs. For example, the first user may increase or decrease the detection sensitivity (or difficulty) by adjusting the threshold for pronunciation errors of the second user downward or upward based on the level to which the second user's first characteristic value belongs.
[0329] For example, if the threshold for current pronunciation errors is set to 0.5 and the second user's first characteristic value is 0.55, the first user (or server (100)) can raise the threshold for the second user's pronunciation errors to 0.6. As a result, the second user can proceed to a lower difficulty level in the next reading, so that they do not lose interest in learning.
[0330] For example, if the current threshold for pronunciation errors is set to 0.5 and the second user's first characteristic value is 0.45, the first user (or server (100)) can raise the threshold for the second user's pronunciation errors to 0.4. As a result, the second user can proceed to a higher difficulty level in the next reading, thereby improving their reading skills without losing interest in learning.
[0331] In addition to pronunciation errors, intonation errors, reading interruption errors, non-response errors, reading speed errors, and concentration loss errors are also divided into levels 1 to n (e.g., levels 1 to 5), and the first user can lower or raise the threshold value for each error type based on the level to which each of the second user's characteristic values belongs.
[0332] For example, the above-mentioned preset threshold can be set as a differentiated value for each age group based on the learner's age information.
[0333] For example, the above age groups may be divided into the first to the nth age groups, the first age group may be the age corresponding to the first and second grades of elementary school, the second age group may be the age corresponding to the third and fourth grades of elementary school, and the third age group may be the age corresponding to the fifth and sixth grades of elementary school.
[0334] For example, the above-mentioned preset threshold can be set to different values depending on the type of reading error.
[0335] For example, the above types of reading errors may include pronunciation errors, intonation errors, reading pause errors, non-response errors, reading speed errors, and concentration loss errors.
[0336] For example, the above-mentioned preset threshold is set by taking into account the age group of the learner and the type of reading error, and the threshold for pronunciation errors in the first age group is 0.8, the threshold for pronunciation errors in the second age group is 0.5, and the threshold for pronunciation errors in the third age group is 0.3.
[0337] For example, the threshold for intonation errors, reading interruption errors, non-response errors, reading speed errors, and concentration loss errors for the first age group is 0.7, the threshold for pronunciation errors, non-response errors, and concentration loss errors for the third age group is 0.4, and the threshold for reading speed errors for the third age group may be 0.3.
[0338] For example, the above-mentioned preset reference value may be automatically determined by an algorithm pre-entered into the server, or may be manually adjusted by the first user through the first terminal.
[0339] For example, when the first user opens a reading session, a user interface for adjusting the threshold value may be provided to the first terminal.
[0340] For example, the above-mentioned preset standard value may be subdivided into multiple level ranges, and the multiple level ranges may be the same as the above-mentioned level 1 to n levels (e.g., level 1 to n levels). For example, the relay reading education method may further include the step of determining whether a characteristic value based on the result of the learner's reading performance falls into any one of the multiple level ranges, and the step of adjusting the standard value to be applied to the learner's next reading based on the level range to which the characteristic value belongs.
[0341] For example, the step of adjusting the reference value may lower the detection sensitivity by adjusting the reference value upward when the characteristic value is higher than the currently set reference value, and increase the detection sensitivity by adjusting the reference value downward when the characteristic value is lower than the currently set reference value.
[0342] For example, the adjustment of the above thresholds can be performed independently for each of the pronunciation error, intonation error, reading interruption error, non-response error, reading speed error, and concentration loss error.
[0343] For example, the step of adjusting the above threshold value may adjust the threshold value upward to 0.6 when the threshold value for the current pronunciation error is set to 0.5 and the learner's characteristic value is 0.55, or adjust the threshold value downward to 0.4 when the threshold value for the current pronunciation error is set to 0.5 and the learner's characteristic value is 0.45.
[0344] For example, the above server can create and manage individualized threshold profiles for each learner.
[0345] For example, the individualized threshold profile mentioned above can be generated by reflecting the learner's past reading performance history, learning progress, and individual weakness information.
[0346] For example, the server may analyze the trend of a learner's level change to evaluate learning progress and provide the evaluation result to the first terminal.
[0347] As such, setting differentiated standards by age group provides the following educational benefits. First, it reduces the dropout rate by providing an appropriate level of difficulty suited to developmental stages. Second, it maintains learning motivation by applying lenient standards to beginner learners, while stimulating a sense of challenge by applying strict standards to advanced learners. Third, by reflecting the cognitive developmental characteristics of each age group, it enables natural skill improvement without an excessive learning burden.
[0348] Figure 11 is a flowchart schematically illustrating a method for classifying reading error types, including providing feedback information.
[0349] Referring to FIG. 11, the reading error type classification method may further include the step (S370) of transmitting feedback information corresponding to the output reading error type to the user terminal.
[0350] If the type of reading error output above is a pronunciation error, the feedback information may include standard pronunciation voice data, pronunciation correction tip text, or voice waveform comparison guidelines.
[0351] For example, the above standard pronunciation voice data may be voice data in which a professional reader pronounces the above standard text or voice data synthesized by a Text-to-Speech (TTS) model from the above standard text.
[0352] For example, the above pronunciation correction tip text may be in the form of, 'Adjust the position of your tongue so that the 'ㄱ' sound does not sound strong,' or 'Lightly touch the tip of your tongue to the roof of your mouth to pronounce the final consonant 'ㄴ' accurately.
[0353] For example, the above voice waveform comparison guideline may be a visual display comparing a user voice waveform and a standard voice waveform side by side.
[0354] The user voice waveform may be generated in the form of a waveform graph expressing amplitude values along the time axis by sampling the preprocessed voice data, and the standard voice waveform may be generated by waveformizing a standard spoken voice synthesized by a Text-to-Speech (TTS) model from the reference text in the same way.
[0355] If the type of reading error output above is an intonation error, the feedback information may include standard intonation voice data or pitch curve information for a reference intonation pattern.
[0356] A standard intonation pattern may refer to data analyzing changes in pitch and intonation when a voice expert recites a specific sentence.
[0357] For example, standard intonation voice data may be voice data generated by a voice expert recording the sentence 'I go to school' according to a standard intonation pattern, or by a Text-to-Speech (TTS) model synthesizing it.
[0358] For example, pitch curve information for a reference intonation pattern may be information configured to allow visual comparison of intonation differences by displaying the pitch curve generated by analyzing changes in the fundamental frequency of a user's speech and the pitch curve of the reference intonation pattern on the same coordinate axis.
[0359] If the type of reading error output above is a reading interruption error, the feedback information may include a re-reading instruction message, information on the number of interruptions, or information on highlighting the interruption point.
[0360] For example, the above re-reading instruction message may be in the form of 'Reading has been interrupted. Please resume reading.'
[0361] For example, the above information on the number of interruptions may be information indicating the number of times the user stopped reading (e.g., '3 times').
[0362] For example, the above break point highlighting information may be UI display information that specifies the location of the sentence where the reading was interrupted to be indicated by visual effects such as changing the text color, underlining, highlighting, or background color emphasis.
[0363] If the type of reading error output above is a non-response error, the feedback information may include a warning message, animation information for prompting reading, or non-response time information.
[0364] For example, the above warning message may be in the form of 'No voice detected. Please start reading.'
[0365] For example, the animation information for inducing reading above may be information that specifies an animation effect causing the microphone icon to blink periodically.
[0366] For example, the no-response time information may be information indicating the user's no-response time as a number (e.g., '5 seconds').
[0367] If the type of reading error output above is a reading speed error, the feedback information may include speed correction guide voice data or a reading speed graph.
[0368] The above speed correction guide voice data may be voice data that induces the user to read aloud at a previously preset standard reading speed.
[0369] For example, the above speed correction guide voice data may be in the form of voice data such as 'Please read a little slower'.
[0370] For example, the reading speed graph above may be graph information that visually compares the user's reading speed with the standard reading speed above.
[0371] If the type of reading error output above is a concentration decrease error, the feedback information may include animation information or encouragement messages to induce concentration improvement.
[0372] The animation information for inducing concentration improvement mentioned above may be information that specifies a dynamic effect providing visual rewards to help the user focus on reading aloud.
[0373] For example, the animation information for inducing concentration improvement mentioned above may be information that fills up star icons at the top of the screen one by one when reading aloud for a certain period of time or longer, or displays an animation of a certain character gradually growing.
[0374] For example, the above encouragement message could be in the form of, "Please read with a little more concentration."
[0375] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0376] The scope of the present invention is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention. Explanation of the symbols
[0377] 100: Server 101: Processor 102: Memory 103: Communication Module 110: First terminal 120: Terminal 2-1 130: Terminal 2-2
Claims
Claim 1 A relay reading education method performed by a server comprises: receiving a request to open a reading group from a first terminal—wherein the reading group includes a first terminal acting as a teacher and a second-1 terminal and a second-2 terminal acting as learners—providing an e-book text to the reading group; granting voice input permission to the second-1 terminal; receiving first reading data from the second-1 terminal and generating first error information including the number of reading errors and time limit exceeded information based on the first reading data; and transmitting first visual information to the second-1 terminal to visually display the reading status based on the first error information. A relay reading education method comprising the step of granting voice input authority to the second-2 terminal or the first terminal based on the first error information, wherein the first error information includes the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit, and the step of transmitting the first time information further comprises the step of generating user interface display information based on the first error information, wherein the user interface display information includes at least one of: i) information instructing that some of a plurality of pre-displayed segments disappear as the number of errors increases; ii) information instructing that the display state of a pre-displayed time progress bar change according to the elapsed time; and iii) information instructing that the display form of a pre-displayed icon change according to an error level - the error level is calculated based on the number of errors and the elapsed time and is a normalized value between 0 and 1 - and the error level is calculated by the following mathematical formula 1. [Mathematical Formula 1] Error level = α × (Number of errors / Maximum allowed number of errors) + β × (Elapsed time / Timeout) where α and β are weights (0 ≤ α, β ≤ 1, α + β = 1), and the maximum allowed number of errors is a pre-set value. Claim 2 ◈Claim 2 was abandoned upon payment of the registration fee.◈ In Claim 1, the step of providing an e-book text to the reading group comprises: receiving a request to provide an e-book text from the first terminal; providing a pre-stored list of e-books to the first terminal; receiving information about an e-book selected from the list of e-books from the first terminal; and dividing the entire text constituting the selected e-book into sentence units, and then sequentially providing a pre-set number of sentences to the reading group, a relay reading education method. Claim 3 ◈Claim 3 was abandoned upon payment of the registration fee.◈ In claim 1, the step of generating the first error information comprises: receiving the first reading data from the 2-1 terminal; applying the first reading data to a pre-trained speech recognition model to convert it into a first speech recognition text; and generating the first error information based on the first speech recognition text and the e-book text provided to the reading group, a relay reading education method. Claim 4 ◈Claim 4 was abandoned upon payment of the registration fee.◈ A relay reading education method comprising: a step of generating a first corrected text by inputting an e-book text provided to the reading meeting into an error pattern backtracking model; -the error pattern backtracking model is a model that receives a correct text, a database of past learner errors, and a current learning context vector as input and outputs a set of predicted error patterns-; a step of generating first synthetic voice data by performing Text-to-Speech (TTS) on the first corrected text; a step of transmitting the first synthetic voice data to the first terminal and the second-1 terminal; a step of generating second error information based on the first synthetic voice data; and a step of granting voice input authority to the first terminal or the second-1 terminal based on the second error information. Claim 5 ◈Claim 5 was abandoned upon payment of the registration fee.◈ In claim 1, the step of transmitting the first time information comprises: the step of transmitting the user interface display information to the 2-1 terminal; and the step of controlling the state of at least one of the graphic elements constituting the user interface displayed on the 2-1 terminal to be changed according to the user interface display information, a relay reading education method. Claim 6 ◈Claim 6 was abandoned upon payment of the registration fee.◈ In claim 4, the error pattern backtracking model searches for candidate errors corresponding to the learner's age group from an error pattern database, calculates an educational value score including a frequency score, a learning transfer score, and a difficulty suitability score for each candidate error, and selects errors whose educational value score is greater than or equal to a preset threshold to generate the first corrected text, a relay reading education method. Claim 7 ◈Claim 7 was abandoned upon payment of the registration fee.◈ In claim 4, the above-mentioned first modified text is a text generated by at least one text generation method among a method of randomly changing the e-book text provided to the above-mentioned reading meeting by a preset ratio or a method of additionally inserting arbitrary characters into the e-book text provided to the above-mentioned reading meeting, a relay reading education method. Claim 8 In a server that performs operations necessary for relay reading education, the server receives a request to open a reading group from a first terminal - the reading group includes a first terminal acting as a teacher and a second-1 terminal and a second-2 terminal acting as learners - provides an e-book text to the reading group, grants voice input authority to the second-1 terminal, receives first reading data from the second-1 terminal, generates first error information including the number of reading errors and time limit exceeded information based on the first reading data, transmits first visual information to the second-1 terminal to visually display the reading status based on the first error information, grants voice input authority to the second-2 terminal or the first terminal based on the first error information, the first error information includes the number of errors resulting from the reading performance and the elapsed time exceeding a preset time limit, generates user interface display information based on the first error information when transmitting the first visual information, and the user interface display information comprises: i) a plurality of preset displayed as the number of errors increases A server comprising at least one of the following: information instructing that some of the segments be extinguished; ii) information instructing that the display state of a pre-displayed time progress bar be changed according to the elapsed time; and iii) information instructing that the display form of a pre-displayed icon be changed according to the error level, wherein the error level is calculated based on the number of errors and the elapsed time and is a normalized value between 0 and 1, calculated by the following mathematical formula 1. [Mathematical Formula 1] Error Level = α × (Number of Errors / Maximum Allowable Number of Errors) + β × (Elapsed Time / Time Limit) where α and β are weights (0 ≤ α, β ≤ 1, α + β = 1), and the maximum allowable number of errors is a pre-set value. Claim 9 ◈Claim 9 was abandoned upon payment of the registration fee.◈ In claim 8, the server receives a request to provide e-book text from the first terminal, provides a pre-stored list of e-books to the first terminal, receives information about an e-book selected from the list of e-books from the first terminal, divides the entire text constituting the selected e-book into sentence units, and then sequentially provides a pre-set number of sentences to the reading group. Claim 10 ◈Claim 10 was abandoned upon payment of the registration fee.◈ In claim 8, the server receives the first reading data from the 2-1 terminal, applies the first reading data to a pre-trained speech recognition model to convert it into a first speech recognition text, and generates the first error information based on the first speech recognition text and the e-book text provided to the reading group. Claim 11 ◈Claim 11 was abandoned upon payment of the registration fee.◈ In claim 8, the server generates a first corrected text by inputting the e-book text provided to the reading meeting into an error pattern backtracking model, -the error pattern backtracking model is a model that receives the correct text, a past learner error database, and a current learning context vector as input and outputs a set of predicted error patterns- generates a first synthesized speech data by performing Text-to-Speech (TTS) on the first corrected text, transmits the first synthesized speech data to the first terminal and the second-1 terminal, generates second error information based on the first synthesized speech data, and grants voice input authority to the first terminal or the second-1 terminal based on the second error information. Claim 12 ◈Claim 12 was abandoned upon payment of the registration fee.◈ In claim 8, the server transmits the user interface display information to the 2-1 terminal and controls the state of at least one of the graphic elements constituting the user interface displayed on the 2-1 terminal to be changed according to the user interface display information. Claim 13 ◈Claim 13 was abandoned upon payment of the registration fee.◈ In Claim 11, the error pattern backtracking model searches for candidate errors corresponding to the learner's age group from an error pattern database, calculates an educational value score including a frequency score, a learning transfer score, and a difficulty suitability score for each candidate error, and selects errors whose educational value score is greater than or equal to a preset threshold to generate the first corrected text, a server. Claim 14 ◈Claim 14 was abandoned upon payment of the registration fee.◈ In Claim 11, the above-mentioned first modified text is a text generated by at least one text generation method among a method of randomly changing the e-book text provided to the reading meeting by a preset ratio or a method of additionally inserting arbitrary characters into the e-book text provided to the reading meeting, a server.