Method and apparatus for generating differentiated oral prompts for learning
By generating personalized verbal prompts and utilizing a conditional language generation model combined with diagnostic and suggestion vectors, the problem of AI coaches lacking personalized communication is solved, thereby improving learning efficiency and interest.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- AGENCY FOR SCI TECH & RES
- Filing Date
- 2024-03-27
- Publication Date
- 2026-04-14
AI Technical Summary
Existing AI coaches lack personalized communication and cannot assess students' learning progress in real time, resulting in low learning efficiency. Furthermore, the information they provide lacks direct relevance, leading to a dull and frustrating learning process.
By generating personalized verbal cues, and utilizing a conditional language generation model that combines diagnostic and suggestion vectors, customized verbal feedback is produced, including emotional settings and adjustments to tone, rhythm, and volume, providing encouragement and guidance.
It improves the personalization and effectiveness of learning, reduces students' frustration, and increases their interest and efficiency in learning.
Smart Images

Figure 2026511871000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing, and more particularly, to a method and apparatus for converting text into speech. More specifically, the present invention relates to a method and apparatus for generating differentiated oral prompts for learning.
Background Art
[0002] Artificial intelligence (AI) tutors available today generally are based on behaviorism in education. Students using such AI tutors are expected to learn through repetition and reinforcement, and the AI tutors provide quick and consistent feedback (such as marks / scores) to inform the students whether what they are doing is correct or incorrect in response to the presented queries. However, such an approach may ignore the identity and personality of the students. Therefore, AI tutors are not suitable for real-time evaluation of the students' learning process.
[0003] Generally, AI tutors lack personalized communication. As a result, students may feel unmotivated because they do not seem to care whether their AI tutors succeed or fail (compared to human tutors). This lack of motivation may cause students to lose focus on learning. If the emotional dimension of learning is not addressed, the learning efficiency of students will be much lower.
[0004] To overcome the above problems, some AI tutors provide learners with a wide range of help information (text / audio / video). However, such a wide range of information lacks direct relevance to the learners' situations. Learners have to struggle to understand the provided additional information, break it down, and then try to identify the parts that can help them. Providing a large amount of information to learners often makes AI-assisted learning boring and rather frustrating.
Summary of the Invention
[0005] The present invention is defined in the independent claims. Several optimal features of the present invention are defined in the independent claims. [Brief explanation of the drawing]
[0006] In the drawings, unless otherwise specified, the same reference numerals generally refer to the same parts through different drawings. Embodiments of the present invention will be better understood and readily apparent to those skilled in the art from the following description in conjunction with the drawings, merely as examples. [Figure 1] Figure 1 shows an overview of an AI-assisted learning method according to an example of the present invention. [Figure 2] Figure 2 shows a framework of an AI-assisted learning method for generating differentiated correction prompts, according to an example of the present disclosure. [Figure 3] Figure 3 shows an example of differentiated correction prompts generated in response to learner input on a learning task. [Figure 4] Figure 4 illustrates three different use cases for generating differentiated correction prompts in response to learner input to a learning task, as an example of the present disclosure. [Figure 5] Figure 5 illustrates a workflow in which differentiated correction prompts are generated in response to learner input for a learning task, as shown in one example of this disclosure. [Figure 6] Figure 6 illustrates how enhanced speech is generated based on a differentiated correction prompt, according to an example of the present disclosure. [Figure 7] Figure 7 illustrates the system architecture of a learner's device or apparatus according to an example of this disclosure. [Modes for carrying out the invention]
[0007] Artificial intelligence-assisted learning (AI-assisted learning) can be improved by providing customized verbal prompts in response to input from a user (or learner) interacting with the AI-assisted learning platform. An example of such improvement is provided in this disclosure.
[0008] In face-to-face classes, human tutors do not always use scores / marks or large amounts of raw information to guide students. Instead, human tutors provide concise verbal prompts to specific learners to help them arrive at the correct answer, to pinpoint exactly where errors occurred, and / or to direct the learner's attention to relevant content. Such verbal prompts are ubiquitous in traditional education and have proven highly effective in face-to-face teaching practices. Creating an AI tutor capable of providing such verbal prompts is challenging. Verbal prompts in existing AI-assisted learning environments lack expressiveness and are generally standardized for all students.
[0009] The examples in this disclosure propose solutions for providing AI tutors that interact with individual learners in a more effective way. By providing human-like feedback, learning is improved for students.
[0010] It has been observed that students need motivation beyond evaluation results (scores and numbers) to progress in their learning. An AI tutor, as illustrated in this disclosure, is configured to provide such motivation. For example, if a student's score is insufficient, the AI tutor is configured to share the evaluation results of the learning task with the student and to provide encouraging words and discoverative voice commands that may act to trigger the student's understanding and learning.
[0011] An AI tutor according to one example of the present invention may take the form of a device for facilitating student learning (i.e., a learner device). The learner device may be a computing device or a mobile device, such as a desktop computer, a laptop computer, a smartphone, a tablet device, and other handheld devices. The learner device may have elements of the device 702 shown in Figure 7. Any image processor, controller, or processor referred to in this disclosure may also have the same elements as those described and shown with respect to the device 702. One or more of the devices 702 may be able to communicate through a communication network 708 such as a wired or wireless network (e.g., WiFi), a mobile telecommunications network, etc.
[0012] The device 702 may be a computing device and comprises several individual components, including but not limited to a processing unit (or processor) 716 and memory 718 for loading executable instructions 720 (e.g., volatile memory such as random access memory (RAM)), where the executable instructions define the functionality that the device 702 performs under the control of the processing unit 716. The device 702 also comprises a network module 725 that enables the device to communicate over a communication network 708 (e.g., the Internet). A user interface 724 is provided for user interaction and may comprise conventional computing peripheral devices such as a display monitor, mouse, and computer keyboard. The device 702 may also comprise a database 726. It should be understood that the database 726 does not have to be local to the device 702. The database 726 may be a cloud database that holds data used to evaluate the learning progress of learners.
[0013] The processing unit 716 is connected via an input / output (I / O) interface 722 to input / output devices (not shown) such as a computer mouse, keyboard / keypad, display, headphones or microphone, and video camera. The components of the processing unit 716 typically communicate via an interconnected bus (not shown in Figure 7) in a manner known to those skilled in the art.
[0014] The processing unit 716 may be connected to a network 708, for example, the Internet, via a suitable transceiver device (i.e., a network interface) or a suitable wireless transceiver, enabling access to other network systems such as the Internet or a wired local area network (LAN) or wide area network (WAN). The processing unit 716 of the device 702 may be connected to one or more external wireless communication-enabled remote servers 704 and / or other learner devices 706 via their respective communication links 710, 712, and 714, via a suitable wireless transceiver device, for example, a WiFi transceiver, a Bluetooth® module, a mobile telecommunications transceiver suitable for Global System for Mobile Communication (GSM®), 3G, 3.5G, 4G, 5G telecommunications systems, etc.
[0015] Instead of the system architecture described above for the computing device 702, the learner device (including the learner device 706) may be a mobile device having the system architecture of the remote server 704, such as a smartphone, tablet device, and other handheld device. Furthermore, one or more learner devices 706 may communicate through other communication networks such as wired or wireless networks and mobile telecommunications networks.
[0016] The remote server 704 may comprise several individual components, including but not limited to a microprocessor 728 and memory 730 (e.g., volatile memory such as RAM) for loading executable instructions 732, the executable instructions defining the functionality that the remote server 704 performs under the control of the processor 728. The remote server 704 also comprises a network module (not shown in Figure 7) that enables the remote server 704 to communicate via a communication network 708. A user interface 736 is provided for user interaction and control, which may take the form of a touch panel display or a keypad, as is common in many smartphones and other handheld devices. The remote server 704 may also connect to a database (not shown in Figure 7) that is not local to the remote server 704 but may be a cloud database. The remote server 704 may also include several other input / output (I / O) interfaces, which may be for connecting to headphones or microphones or speakers (audio devices), subscriber identification module (SIM) cards, flash memory cards, USB-based devices, etc.
[0017] The software and one or more computer programs installed on devices 702, 704, and / or 706 may include one or more software applications for instant messaging platforms, audio / video playback, internet accessibility, operation of device 702, remote server 704, and / or learner device 706 (i.e., operating systems), network security, file accessibility, database management, and the like. The software and one or more computer programs may be encoded on a data storage medium such as a CD-ROM, on a flash memory carrier, or on a hard disk drive and supplied to a user of device 702, remote server 704, or learner device 706, and read using a corresponding data storage medium drive, such as a data storage device (not shown in Figure 7). Such application programs may be downloaded from network 708. The application programs are read by processing unit 716 or microprocessor 728, and their execution is controlled. Intermediate storage of program data may be achieved using RAM 720 or 730.
[0018] Furthermore, one or more steps of a computer program or software may be executed in parallel rather than sequentially. One or more computer programs may be stored on any machine or computer-readable medium that may be inherently non-temporary. Computer-readable mediums may include storage devices such as magnetic or optical disks, memory chips, or other storage devices suitable for interfacing with a general-purpose computer or mobile device. Machine or computer-readable mediums may also include hardwired media, such as those exemplified in internet systems, or wireless media, such as those exemplified in wireless LAN (WLAN) systems. When a computer program is loaded onto such a general-purpose computer and executed, it effectively provides an apparatus for carrying out the steps of the computing method in the examples described herein.
[0019] In one example, the learner device is configured to determine the learner's identity based on password-based authentication, multi-factor authentication, biometric authentication (such as facial and fingerprint recognition), etc., before granting access to learning tasks. After access is granted, the learner device is configured to allow the learner to access questions / tasks created to fulfill learning objectives set for (or set by) the learner. The learner may be permitted to access a history of questions / tasks attempted by the learner.
[0020] A learner device may be configured to generate, based on the input provided by the learner in response to each question (or learning task), a concise, relevant, and differentiated instruction (or differentiated response) that may be audio, video, and / or text feedback. One example of a differentiated response is praise given to the learner to correct an answer. This praise may be a differentiated oral prompt, which may include a modified or manipulated voice that includes acoustic / audio cues to emphasize and / or deem emphasize parts of the voice, for example, so that the praise produces a more personalized sound. Another example of a differentiated response is a differentiated correction prompt (or, for short, a “correction prompt”) to guide the learner to the correct answer when the learner’s input is deemed an inaccurate, incomplete, or insufficient response to the learning task. For example, a differentiated correction prompt may be a differentiated oral prompt, which may include a modified or manipulated voice that includes acoustic / audio cues to emphasize and / or deem emphasize parts of the voice read aloud to the learner. A differentiated oral prompt may include visual cues, which will be detailed later.
[0021] For clarity, in this disclosure, a differentiated instruction (or differentiated response) refers to a response to a learning task attempted by a learner, which is output by a learner device regardless of whether the response is considered the correct response to the learning task. A corrective prompt or differentiated corrective prompt or differentiated assistance prompt refers to a response to a learning task attempted by a learner, which is output by a learner device when the response is considered an incorrect response to the learning task. As explained in the previous paragraph, a differentiated corrective prompt is a differentiated instruction (or differentiated response), but a differentiated instruction or response may not be a differentiated corrective prompt. In an example of this disclosure, a differentiated response results in the generation of an oral prompt or differentiated oral prompt that refers to an audio output in response to feedback (or answer) provided by a learner. The term "differentiated" means "adjusted", i.e., a differentiated oral prompt is an oral prompt adjusted to the user. The final form of the oral prompt or differentiated oral prompt to be output should include audio that has been modified or manipulated to include an acoustic / audio cue to emphasize and / or de-emphasize a portion of the text or audio. Before the oral prompt or differentiated oral prompt is modified or manipulated to include an audio cue, it can be referred to simply as an oral prompt or differentiated oral prompt for simplicity.
[0022] In some examples, a differentiated oral prompt may exclude reading the oral prompt aloud to the learner via an audio device and include only text that has been modified or manipulated to include visual cues generated based on an acoustic / audio cue. Such visual cues emphasize and / or de-emphasize portions of the text that are displayed to the learner according to the acoustic / audio cue. Similarly, before the differentiated oral prompt (or oral prompt) is modified or manipulated to include such visual cues, it can be referred to simply as an oral prompt or differentiated oral prompt for simplicity.
[0023] The ability of the learner device to generate concise, relevant, and differentiated correction prompts for the learner for learning tasks helps the learner recognize their mistakes before losing concentration, especially for those with short attention spans. The feedback provided to the learner in traditional AI tutors is generally not concise and may cause delays in finding the cause of the error. To summarize inaccurate feedback, the learner may need to perform extra work, resulting in loss of concentration, interest, and / or motivation, and hindering learning efforts.
[0024] By providing direct and / or concise feedback using differentiated correction prompts, the learner device of this example can reduce the learner's frustration and precious time, thereby enabling the learner (or student) to progress faster.
[0025] Furthermore, the differentiated instructions (or differentiated responses) that can be provided by the learner device are configured to resonate with the student. This can be achieved by including, as will be described later with reference to FIGS. 2 and 3, a voice intensifier. The voice intensifier is designed to convey meaning to the learner through explicit meaning and / or implicit meaning, thereby increasing the learner's level of learning engagement.
[0026] An example of a method for generating differentiated instructions according to the diverse learning abilities of the student, which can be executed by the learner device, is discussed in the following paragraphs.
[0027] Referring to Figure 1, a conditional language generation model is used to generate differentiated instructions according to a student's learning task. The inputs to the model are data from the learning task 102 and the evaluation results 104 for this learning task 102. The output of the model is differentiated oral prompts regarding the student's performance on the learning task 102. The differentiated oral prompts may include help or guidance information to facilitate learning and / or to provide praise / encouragement.
[0028] On the input side, a diagnostic vector 110 derived or obtained from an AI assessment module can be used to represent each student's performance 104 on the learning task 102. Such an AI assessment module may operate based on a neural network or a predefined algorithm. The diagnostic vector may be a vector having one or more parameters related to a student's performance on the learning task 102. One of the parameters may be the score given for completing the learning task 102. In addition to considering the performance (i.e., score) 104 obtained for completing the learning task 102, the diagnostic vector may also take into account each student's academic background 106 and teacher evaluation 108. In some examples, the diagnostic vector may also include the student's answers as one of its parameters. In one example, one of the parameters may be the student's academic ranking in the class or school for the learning task, a review given to the student by a peer or educator for the learning task, a comparison value indicating the student's performance compared to a previously completed task, compared to another student, or compared to the average performance of a group of students. Mathematically, the diagnostic vector is,<a,b,c> It may also be expressed in the form of a, b, c, and , for example, a being the score obtained for completing the learning task, b being the student's ranking in a group of students who complete the same learning task, and c being a pointer to a string of text related to the teacher's evaluation. The diagnostic vector 110 can accurately reflect the learning problems the student is facing with respect to the learning task 102 at the current stage.
[0029] As shown in Figure 1, the diagnostic vector 110 obtained from the AI evaluation module is input to the conditional language generation model 115. The diagnostic vector 110 may include information about how well the student performed, the student's response, and / or information to determine which correction prompts to generate and communicate to students who have completed the learning task 102. The student's response may be incorporated together with the diagnostic vector 110, or it may be input separately to the AI evaluation module.
[0030] In the example in Figure 1, the diagnostic vector 110 contains all the information necessary for the language model 115 to generate conditional feedback on the student's learning. This includes the learning task 102, grades 104, academic record 106, and teacher evaluation 108, which may take the form of the following vector.
[0031] 1. The learning task vector 102 may be a one-hot or distributed vector pre-assigned to each learning task. This may be one of the inputs to the evaluation module (i.e., part of a specific task 202 in Figure 2, described later). 2. The performance vector 104 may be the output of the evaluation module. For different learning tasks, this vector may represent the result of a qualitative or quantitative evaluation of the learner's input. 3. The history vector 106 may be the aggregated grades of past students. 4. The evaluation vector 108 is the teacher's input. Such an evaluation vector is arbitrary. If there is no input from any human tutor, this vector can be set as the average of the teacher's inputs for all training or learning data (i.e., data on the student's learning).
[0032] The output of a typical language generation model is text only. However, text itself is ambiguous in conveying paralinguistic cues such as tone, pitch, stress (or prosody), volume, and speed, which are important for learning. To avoid losing these important cues, the method may include a step of deriving an suggestion vector 118. Such an suggestion vector may include parameters relating to tone, pitch, stress (or prosody), volume, and speed. For example, the suggestion vector may include values indicating levels of speech tone, pitch, stress, volume, and speed. The suggestion vector 118 is computed together with a text command 116. The text command 116 indicates a word that should be spoken orally, for example, through a speaker device. The text command 116 may be generated based on information provided in the diagnostic vector and / or the learner's answer or response to the learning task 102. For example, the text command may be predetermined feedback given based on the learner's answer or response to the learning task 102. The suggestion vector 118 indicates how the word should be spoken by emphasizing paralinguistic cues during the generation of oral prompts to the learner. The suggestion vector 118 may be generated based on the information provided in the diagnostic vector, the learner's answer or response to the learning task 102, and / or based on the text instruction 116.
[0033] For example, if a learner fails to complete a learning task or does not perform it sufficiently well, a differentiated corrective prompt in the form of an oral prompt, to be presented by the learner device, may indicate the score obtained, provide the correct answer, provide an explanation of the answer, explain what caused the failure, inform the learner how they compare to other peers, and / or provide teacher or peer feedback. The correct answer provided in case of failure may include vocal emphasis, which may be auditory and / or visual. If the learner successfully completes the learning task, a differentiated response in the form of an oral prompt may indicate the score, provide correctness, praise, how they rank compared to other peers, and / or indicate what was done well.
[0034] Furthermore, the method for generating differentiated oral prompts comprises the step of processing an implied vector 118 together with a text command 116. Specifically, the text command 116 and the implied vector 118 are sent to a text-to-speech (TTS) synthesizer 120 comprising a text analysis module 125 and a digital signal processing (DSP) module 130. The text analysis module 125 generates an audio transcript of the text to be read, provided by the text command 116, dividing and marking the text into prosodic units such as phrases and sentences. If the text to be read contains numbers and abbreviations, they should be converted into their written word equivalents. The generated audio transcript or transcription is then sent to the DSP module 130 along with prosodic information (e.g., desired intonation and rhythm). The prosodic information is provided by or derived from the implied vector 118. The DSP module 130 then generates a synthesized speech corresponding to the text to be read.
[0035] Unlike conventional TTS synthesizers that produce a uniform style of speech that can only express the superficial meaning of the spoken text, the TTS synthesizer 120 is configured to modify the prosodic information before the text is read aloud to the learner. The modifications made to the prosodic information are provided in the form of suggestion vectors. Suggestion vectors may be generated considering diagnostic vectors associated with each learning task and / or the learner's response or answer to the learning task. Diagnostic vectors may indicate the learner's performance and may include the learner's response or answer to the learning task for consideration. The generated differentiated oral prompts emphasize specific parts of the speech according to the prosodic information, and as a result, the learner learns better compared to conventional uniform style speech. In particular, the TTS synthesizer 120 is configured to manipulate finer areas of the speech signal using acoustic variations in intonation, temporal rhythm, and volume to subtly attract attention.
[0036] For example, if the diagnostic vector indicates that the learner failed to complete the learning task, the suggestive vector may be configured to make the verbal prompt to be generated sound more empathetic, more encouraging, and more instructive. The verbal prompt may also be stressed in parts to point out the learner's mistake. If the diagnostic vector indicates that the learner successfully completed the learning task, the suggestive vector may be configured to praise the verbal prompt to be generated.
[0037] Expressive (emphasized) oral prompts can be used to highlight hints for successfully completing learning tasks. This can give learners a new level of understanding. For example, certain parts of oral prompts can be made louder and slower to provide learners with acoustic cues so that they can easily identify "hidden" meanings between texts. The ability to generate customized speech for different learners can also infuse AI education with attention and authenticity, and thus encourage learners to put in more effort along the learning journey.
[0038] In another example, text command 116 could be feedback generated with annotation words / sentences. Such annotation words / sentences would include what today's AI tutor's monotonous voice is saying.
[0039] However, unlike today's AI tutors, the suggestion vector 118 in this example can be designed to allow the AI tutor to actively take ownership of the learning process through expressive commands, capture the learner's attention, and stimulate the learner in a differentiated manner. The suggestion vector 118 can be associated with a set of sentiment settings defined for educational purposes. Such sentiment settings are as follows: • Guiding and / or encouraging; Delightful and / or joyful; • Criticism; and ·absolute
[0040] For example, the suggestion vector 118 may contain four values, each associated with the degree or level of each of the four settings described above.
[0041] In addition to emotional setting, the suggestion vector 118 can also indicate errors made by the learner and remind the learner to correct them by emphasizing words and phrases within the text instruction 116 (i.e., suggesting to the learner how his or her answer should be answered, corrected, and / or improved).
[0042] An emphasis mask (or stress mask) (e.g., 238 in Figure 2) may be employed to implement a rhythmic structure in speech, i.e., a differentiated oral prompt generated for the learner to convey the implicit meaning of the text command 116. Such an emphasis mask may be part of the suggestion vector 118, include the suggestion vector 118, or be provided separately from the suggestion vector 210. Several examples are shown below: If a student is struggling to learn, a language model can be implemented to generate a text instruction 116 with a “instruction and / or encouragement” sentiment setting. In this example, the text-to-speech synthesizer 120 and the digital signal processor (DSP) 130 apply a stress mask according to iambic meter (unstressed and stressed rhythms), and such iambic meter tends to have a gentle, flowing quality that gives a sense of encouragement. The speech of the oral prompt is produced with a natural, pleasant rhythm, like a steady heartbeat. Students feel reassured by the sense of stability and solidity. Specifically, iambic meter is a poetic style in which stress is placed on every beat or every syllable. For example, an unstressed (or unstressed) syllable is followed by a stressed (or stressed) syllable. If the learner is easily distracted and their attention is diverted from what they are learning, the "absolute" emotional setting can employ a stress mask following a stressed-stress rhythm. The continuous stress on syllables creates a sense of urgency, which alerts the learner to immediately return their attention to what they are learning. Specifically, stressed-stress is a poetic style in which every beat-syllable is stressed. For example, one stressed (or accented) syllable is followed by another stressed (or accented) syllable. When students are making progress, the "delightful and / or joyful" emotional setting can employ stress masks according to stressed and unstressed rhythms to create an active, upbeat rhythm. Traditionally, writers associate stressed rhythms with feelings of happiness. Specifically, stressed rhythm is a style of poetic poetry that has a "descending rhythm." For example, a stressed (or accented) syllable lies on the first beat followed by an unstressed (or unstressed) syllable. When a learner makes a mistake, the emotion of "blame" can be conveyed by employing a stress mask in a stress-unstressed-unstressed rhythm. This creates a sense of strong momentum, reflecting the tutor's rising anger and passion, and evoking a sense of urgency and excitement. In particular, the stress-unstressed rhythm is a poetic style that has a stressed (or accented) syllable followed by two unstressed (or unstressed) syllables. In addition to the above, emphasis masks can be set to emphasize (stress) specific words in a sentence to tell learners what is important, bringing clarity to the implicit meaning using text instruction 116. For example, the following sentences (sample text instruction 116) have exactly the same meaning. However, by stressing different words, differentiated oral prompts generated in audio or text for a sentence draw the learner's attention to different parts of the sentence. [Table 1]
[0043] The emotional settings and stress mask patterns described above can be summarized in Table 1 below. [Table 2]
[0044] The following paragraphs describe a method for generating expressive speech by emphasizing (highlighting) specific words / phrases in a sentence, which can be performed by the learner device described above.
[0045] Speech is irreplaceable in classroom education, and it can effectively communicate to the target audience in a way that is easily received and meaningful, going beyond text. In such traditional educational settings, heuristic methods and suggestive corrections are widely used as a scaffold to help students (or learners) discover their own errors.
[0046] A method for generating expressive speech that can be performed by a learner device leverages the intuitive and suggestive nature of speech characteristics to provide stronger acoustic cues in selected regions of a speech signal, thereby providing guidance and enabling learners to independently identify errors and formulate corrections.
[0047] In the example of the method for generating expressive speech described below, students provide oral or spoken answers to a specific task 202, which may take the form of a question. The oral or spoken answers are the individual responses 204 of the students.
[0048] Referring to Figure 2, this method comprises two stages 220 and 240. The first stage 220 relates to the generation of clear oral prompt text or text string 236 that highlights a mask 238 based on a specific task 202 and the individual responses 204 of the students. The second stage 240 relates to the generation of emphasized speech 255 for separate oral prompts. The term “emphasized speech” as described herein refers to speech produced by a DSP module (e.g., 130 in Figure 1) and at least a portion of the speech has acoustic changes in intonation, temporal rhythm, and volume. The acoustic changes are configured to subtly attract attention.
[0049] In the first stage 220, an automated evaluation step 222 is provided for the student's individual responses 204, performed by the learner device. The evaluation results are used to derive personalized help prompts (i.e., differentiated correction prompts). During the evaluation step 222, an evaluation is made for the student's individual responses 204 to a particular task 202. For example, if the particular task 202 is a language learning task, the learner device is configured to detect one or more pronunciation problems in the student's speech during the speech evaluation 224. If the particular task 202 is a multi-stage mathematical problem, the learner device is configured to break down the mathematical problem into its solution stages and, during the quantitative reasoning stage 226, to accurately point out possible misperceptions, misunderstandings, and gaps in knowledge in the student's individual responses 204.
[0050] The AI assessment module used to evaluate student responses can operate using speech assessment models, quantitative reasoning models, and / or language generation models. The AI assessment module may also be a neural network, and the model may be trained on annotated datasets for specific learning tasks. As an example, Google's Minerva language model (>100B parameters) demonstrates the ability to process scientific and mathematical questions (or tasks) formed in natural language and generate stepwise solutions. The Minerva model also incorporates prompting and assessment techniques.
[0051] For example, the learning task may be a mathematical task in the form of a multiple-choice question (MCQ). In this example, the AI evaluation module is configured to perform a binary evaluation (i.e., true or false) 228 for each answer choice of the MCQ. The AI evaluation module may use a conditional language generation model 230 that focuses on generating meaningful prompts based on the binary evaluation results 228. The evaluation does not have to focus on identifying the correct answer, but rather on providing meaningful prompts for each answer choice, whether correct or incorrect. When the learning task is a multiple-choice question, it has been observed that having a large amount of training data to train the AI evaluation module can maximize its performance in generating meaningful, differentiated instructions for each MCQ answer. Meaningful prompts may include the correct answer to the MCQ and / or praise if the correct answer is given, as well as corrective prompts that guide the learner to the correct answer if an incorrect answer is given to the MCQ. For example, if the learner selects an incorrect answer to the MCQ, they are given verbal prompts that guide them to the correct answer, rather than verbal prompts that directly give the correct answer.
[0052] The learner device is configured to output parameters for generating differentiated oral prompts from the conditional language generation model 230 (regardless of whether the learner obtained a correct or incorrect answer to the question) based on the evaluation result 228, the parameters including an emphasis mask 238 (which may be part of, include, or be provided separately from, the suggestion vector 118 in Figure 1), and a text command or text string 236 to be read. Such parameters are used in a second stage 240 to provide acoustic (or audio) cues and / or visual cues. The text command or text string 236 is modified by the expression TTS synthesizer 242. At least a portion of the synthesized speech of the text string output by the TTS synthesizer 242 (i.e., acoustic cues or audio cues) has acoustic changes in intonation, temporal rhythm, and volume. The acoustic changes may be modified according to the prosodic information provided by the emphasis mask 238. Important parts of the audio cue are either emphasized (e.g., set to a higher volume) or de-emphasized (e.g., set to a lower volume). The audio emulation unit 244 is provided to output the thus emphasized audio as an audio file or to play it back through an audio speaker.
[0053] The audio cues may be configured to display visual cues for differentiated oral prompts. These visual cues may appear when differentiated oral prompts are read aloud. The visual cues may appear in text format and, according to the corresponding audio cues, visually highlight important parts of the displayed text. For example, the emphasis applied to important parts of the displayed text may take the form of uppercase letters, color, bold, underline, italics, or colored / shaded emphasis. In one example, only visual cues highlighting important parts may be displayed, and no oral prompts to be output through the speaker may be provided.
[0054] Furthermore, the voice intensifier 244 can be configured to acoustically and / or visually highlight (or emphasize) or de-emphasize important portions of audio and / or visual cues at the syllable, word, phrase level, and / or at different portions of a differentiated command.
[0055] Specific examples of the application of the two stages 220 and 240 for generating a differentiated correction prompt and generating an emphasized voice are described as follows with reference to FIGS. 2 and 3.
[0056] In this example, a language learning task (or question) is presented to a student to read a Chinese sentence on a learner device on which software providing the learning task is installed. The language learning task is a question displayed on the screen of the learner device. Through the graphical user interface, the student submits a response (or answer) using the microphone of the learner device. The response to the learning task is received and evaluated by the learner device. The AI evaluation module, which is a component of the software, is activated to evaluate the response. In this example, the student's response to the learning task is an incorrect reading of the Chinese phrase "forty-four stone lions (四十四体の石獅子)". After evaluation, the AI evaluation module instructs the learner device to read out the correct answer with a differentiated oral prompt through the speaker of the learner device. Specifically, the correct answer emphasizes a part 315b of the student's response that was evaluated as incorrect. In this case, the Chinese character "石 (stone)" was misread by the student.
[0057] The AI evaluation module may be a neural network that has received machine training to recognize misreadings, or a non-neural network algorithm configured to recognize misreadings.
[0058] In this example, the text string 236 read by the learner device is the phrase "Forty-four stone lions (Forty-four stone lions)". An audio stream or file is generated containing synthesized speech 310 of the text string 236. Upon recognizing a pronunciation error, the AI evaluation module generates an emphasis mask 238 (which may be part of an suggestion vector, include an suggestion vector, or be provided separately from an suggestion vector) to send to the TTS synthesizer 242. The TTS synthesizer 242 modifies the prosody of the mispronounced character "stone (stone)" in the synthesized speech 310 to be louder and slower according to the emphasis mask 238. The modified synthesized speech 310 constitutes a differentiated oral prompt. When the learner device reads the correct answer, including this differentiated oral prompt, through the speaker, the portion of the correct answer with the modified prosody acts as an acoustic cue to the student to help them understand that the student mispronounced "stone (stone)".
[0059] In this example, a visual representation of the differentiated oral prompts is also provided by the speech intensifier 244. The speech intensifier 244 is configured to box up errors (see box 315a in Figure 3) and to indicate incorrectly pronounced words in a different font color when the correct answer is read aloud through the learner device's speaker. In this way, differentiated correction prompts with multimodal stimuli (including intensified audio and visual guidance) are presented to the learner to highlight (or pinpoint) the errors that need to be corrected. After being presented with differentiated correction prompts, the learner may be asked to retry the learning task. This helps the student learn the correct pronunciation of the words.
[0060] An example of a conditional language generation model (i.e., a prompt generation model) can be summarized by the following conditional entropy for generating differentiated multimodal instructions:
number
[0061] In particular, the learning task, the learner's responses (e.g., answers given by the learner), and task-specific diagnoses (evaluated by the AI evaluation module) are inputs to a conditional language generation model for determining personalized and differentiated help prompts for the learner. The output of the conditional language generation model is a multimodal modified prompt, which comprises a synthesized speech (i.e., a generated prompt with emphasis) read aloud to the learner, and at least a portion of the synthesized speech has acoustic variations in intonation, temporal rhythm, stress, and / or volume.
[0062] Instead of objective evaluation, the learner devices described in the examples of this disclosure are configured to subjectively assess overall performance through two or more application tasks. In one example, a first application task is provided to help a student solve a mathematical (vocabulary) problem, and a second application task is provided to correct a student's pronunciation in language education. Another application task may be provided to help a student have scientific understanding.
[0063] Figure 4 illustrates the steps for generating acoustically and / or visually differentiated verbal prompts, which can be used in many use cases. For example, there may be multiple students attempting to learn a task. For each student, differentiated prompts are generated based on the student's performance on their current learning task (or application task) 402a, 402b, and 402c (i.e., based on diagnostic vectors).
[0064] Student responses may be audio input for language learning task 402a, numerical input for mathematical learning exercise 402b, or text input for scientific comprehension task 402c. In this example, audio input is for the language task, numerical input for the mathematical task, and text input for the scientific task, but it should be noted that such audio input, numerical input, and / or text input are applicable to all of the language, mathematical, and scientific tasks. For example, numerical input may be an answer to a multiple-choice question on any subject, and text input and / or audio input may be provided as an answer to any subject.
[0065] The evaluation process for student responses, i.e., 402a, 402b, and 402c, is task-specific and may involve quantitative inference such as speech evaluation 424a and / or error source detection techniques 424b and fact-checking 424c. Neural networks may be used in the evaluation process, or they may be less complex algorithms for comparing student responses against correct answers stored in a database.
[0066] Following evaluation, concise, relevant, and differentiated oral prompts are generated by the speech enhancer 244 in Figure 2, based on the evaluation results for each learning task. In one example, the learner device is configured to present multimodal differentiated correction prompts when the learner's response is deemed inaccurate for the learning task. For example, selected fine-grained regions of the synthesized speech signal are manipulated to include acoustic changes in intonation, temporal rhythm, and volume, providing stronger acoustic cues to guide the student to make corrections themselves. Visual cues may also be displayed synchronously (or unsynchronously) with the manipulated acoustic cues to provide an aided learning platform.
[0067] Figure 5 illustrates another example of Method 500 for generating one or more differentiated oral prompts for a learner (or student). Such Method 500 is performed on a learner device or apparatus (such as a mobile or desktop computing device) made available to the learner. Method 500 comprises presenting a learning task to the learner on a display and receiving the learner's feedback (or answer to the learning task). The display may be part of the learner device or may be an external display unit (e.g., an LCD screen) of the learner device (such as a desktop). The learning task (or application task) may be a language learning problem, a mathematical problem, a scientific problem, or a problem that can be used to satisfy the learner's learning needs. It should be understood that the learning task may also cover other fields of study such as science, accounting, and computing. The learner's answer to the learning task may be in the form of audio input, numerical input, and / or text input.
[0068] Method 500 comprises, in step 505, evaluating the learner's response to determine a diagnostic input (e.g., a diagnostic vector), and then, in step 510, using a prompt generation model 515 to perform a diagnosis or evaluation based on the diagnostic input to generate a differentiated correction prompt 528 (e.g., in the form of a text string, audio, image, and / or video) and a differentiated speech enhancement mask 522. In this example where the learner's response is incorrect, the differentiated correction prompt 528 includes a text string to be processed.
[0069] The diagnostic or evaluation process of Method 500 is task-specific and may involve speech evaluation and / or quantitative inference, such as error source detection and / or fact-checking. The prompt generation model 515 may include a library of prompts, each prompt being task-specific and / or feedback-specific. Differentiated correction prompts 528 may not be direct answers or explicit solutions to the learning task. Correction prompts 528 may be hints to guide the learner to the correct answer. Providing such correction prompts 528 guides the learner to acquire the skills and knowledge necessary to manage similar learning tasks. Providing the correct answer directly may not achieve this objective.
[0070] In one example, the diagnostic input is a diagnostic vector and may include MCQ answers, correct / incorrect indications, or marks / scores provided by the learner device (after evaluation in step 510). The differentiated emphasis mask 522 is an emphasis mask vector (which may be part of the suggestion vector, contain the suggestion vector, or be provided separately from the suggestion vector) used to emphasize and / or deem a portion of the audio output. The differentiated emphasis mask 522 may include data indicating pitch, volume, and / or tempo settings to be adjusted for the differentiated oral prompt 575 that can be played on the speaker. For example, a character counter may be used to determine the syllables, words, and phrases of the differentiated correction prompt.
[0071] In step 530, the differentiated correction prompt text string 528 and the differentiated speech emphasis mask 522 are input to an AI text-to-speech module (TTS module), which performs local prosodic control to apply the speech emphasis / adjustment specified by the speech emphasis mask 522 to the text string. The emphasis / adjustment may be performed at the syllable, word, letter, and / or phrase level. The output of the TTS module may be a speech signal containing the emphasis / adjustment specified by the speech emphasis mask 522. Such a modified or adjusted speech signal may be an analog signal. In step 560, the output of the TTS module is sent to a digital signal processor to convert the modified or adjusted speech signal into a digital format. In step 535, the output of the TTS module is also sent to an AI speech recognition model for speech recognition to recognize the syllables, words, letters, and / or phrases to be spoken.
[0072] In step 545, temporal information is extracted for these recognized syllables, words, letters, and / or phrases for temporal alignment. The AI speech recognition model 535 may be a model trained by machine learning. The temporal alignment information from step 545, the output from the TTS module, and the differentiated emphasis mask (speech emphasis) 522 are input to a digital signal processor (DSP) for processing in step 560. The DSP processes these inputs to generate a digitized, playable speech signal for the output from the TTS module, which is temporally aligned according to the pitch, tempo, and / or volume information specified in the speech emphasis 522. The output of the DSP is a differentiated oral prompt 575 (in a digital audio file or audio stream format) which can be played back by a media player and made audible through a speaker. The DSP module may also perform enhancements to make the differentiated oral prompt 575 sound better.
[0073] The AI speech recognition step 535, the temporal alignment step 545, and the DSP processing step 560 in Figure 5 can be summarized as the digitization step 540. Figure 6 illustrates an example of the digitization step 540 in Figure 5. Referring to Figure 6, the correction (or revision) prompt 628 and the emphasis mask 622 are generated based on learner-specific and task-specific diagnostic results and provided to the text-to-speech module 625 (corresponding to the TTS module described in relation to Figure 5). The text-to-speech module 625 converts the text string of the correction prompt 628 into an audio signal and adjusts the audio signal to include the emphasis specified by the emphasis mask 622. The adjusted audio signal output from the text-to-speech module 625 is processed by an automatic speech recognition (ASR) module (corresponding to the AI speech recognition model described in relation to Figure 5), which segments the audio signal into a first series of frames 632 using the short-time Fourier transform (STFT) algorithm 627. Each frame corresponds to the spectral characteristics of an individual utterance (i.e., a syllable, word, letter, or phrase). These frames facilitate the temporal alignment of the generated digitized enhanced speech. During temporal alignment, at least one portion 634 of the first set of frames 632 is retained (unmodified), and at least one other portion 636 of the first set of frames 632 is modified based on the enhancement mask 622. The prosodic parameters (pitch shift, tempo stretch, gain, etc.) of at least one modified frame 638 are adjusted for combination with the retained frame 634 to form a second set of frames 640, and then output. The second set of frames 640 is then processed by the inverse short-time Fourier transform (ISTFT) algorithm 645 to produce a finely tuned synthesized speech (DSP-generated audio output) for the learner.
[0074] A digital signal processor (DSP) may be used to perform the Short-Time Fourier Transform (STFT) algorithm 627 and the Inverse Short-Time Fourier Transform (ISTFT) algorithm 645. The DSP may be used to convert the tuned audio signal and the first set of frames 632 output from the text-to-speech module 625 into a digital format for processing. The DSP may also be used to process at least one portion 634 of the first set of frames 632, at least one other portion 636 of the first set of frames 632, at least one modified frame 638, and a second set of frames 640, and output them in digital format.
[0075] Figures 5 and 6 illustrate possible solutions for differentiated oral prompts in learning and should not be interpreted as limiting them. As technology evolves, new (better) solutions for such tasks may exist that apply concepts similar to those covered in this disclosure.
[0076] The following paragraphs provide additional examples of differentiated verbal prompt generation, accompanied by subjective assessments of learning tasks requiring quantitative reasoning.
[0077] In the first example, learners are presented with the following mathematical problem: Example 1: Jack had two apples. He ate one. He plans to buy another one tomorrow morning. How many apples will Jack have tomorrow?
[0078] Subsequently, in response to learner feedback, multimodal correction prompts are presented. In this example, the mathematical problem is a polynomial choice question (MCQ) in which the learner selects an answer from multiple answer choices. As shown below, once the learner selects an answer, a corresponding oral prompt is generated. The diagnostic vector in this example could be the numerically selected answer. With the diagnostic vector as input, the evaluation module can retrieve from the database a text command that shows the text of the oral prompt associated with the selected answer, and a speech emphasis mask associated with the text command. [Table 3]
[0079] For this learning task, if the learner selects option A (i.e., answer 1), the generated differentiated corrected prompt will be "He ate one and plans to buy another tomorrow," and the generated speech emphasis mask will be "He plans to buy another." The oral prompt, when read aloud, emphasizes the words "He plans to buy another" by making them louder and slower compared to the rest of the sentence. Furthermore, the text of the oral prompt to be displayed is adjusted to emphasize "He plans to buy another" by displaying it in a different mode (e.g., bold, color, highlight, underline, italics, etc.) compared to the rest of the sentence. In another example, the words "He plans to buy another" may be read progressively slower and with more variation in intonation, so that the learner can infer from the visual and / or audio cues to perform subtraction (i.e., "2-1") followed by addition, which is necessary to arrive at the correct answer.
[0080] Similarly, if the learner selects option C (i.e., the answer is 3), the differentiated corrected prompt generated would be "He ate one today and plans to buy another tomorrow," and the generated speech emphasis mask would be "He ate one today." The oral prompt, when read aloud, emphasizes the words "He ate one today" by making them louder and slower compared to the rest of the sentence. Furthermore, the text of the oral prompt to be displayed would be adjusted to emphasize "He ate one today" by displaying it in a different mode (e.g., bold, color, highlight, underline, italics, etc.) compared to the rest of the sentence. In another example, the words "He ate one today" may be read progressively slower and with more variation in intonation, so that the learner can infer from the visual and / or audio cues to perform subtraction before addition (i.e., "2+1") necessary to arrive at the correct answer.
[0081] When the learner selects the correct answer (in this case, option B), the generated oral prompt and speech emphasis mask become "Well done." This oral prompt may be at a higher pitch when read aloud to motivate the learner. Modifying the prosody of the oral prompt for the correct answer is optional.
[0082] In the second example, the learner is presented with another mathematical problem, as shown below.
[0083] Example 2: My number has four digits and a 7 in the hundreds place. The highest value digit in my number is 2. The lowest value digit in my number is 6. My number has 3 fewer tens digits than hundreds digits. What is my number?
[0084] After receiving learner feedback (or answers) to a question (learning task), a multimodal correction prompt is presented. In this example, the mathematical problem is a polynomial choice question (MCQ) in which the learner selects an answer from multiple answer choices. In another example, the learner may be asked to provide an answer in response to a learning task as a numerical input (i.e., by keying a four-digit number into the learner device). As shown below, once the learner selects an answer, a corresponding verbal prompt is generated. [Table 4]
[0085] For this learning task, if the learner selects option B (i.e., the answer is 3746), the generated differentiated corrected prompt is "The digit with the highest value is 2," and the generated speech emphasis mask is "The highest value is 2." The oral prompt, when read aloud, emphasizes the words "The highest value is 2" by making them louder and slower compared to the rest of the sentence. Furthermore, the text of the oral prompt to be displayed is adjusted to emphasize "The highest value is 2" by displaying it in a different mode (e.g., bold, color, highlight, underline, italics, etc.) compared to the rest of the sentence. In another example, the words "The highest value is 2" may be read progressively slower and with more variation in intonation, allowing the learner to infer from the visual and / or audio cues that the thousands digit of the four-digit value is 2.
[0086] If the learner selects option C (i.e., the answer is 2736), the differentiated correction prompt generated will be "My number has 3 fewer tens than hundreds" and the generated speech emphasis mask will be "3 fewer tens than hundreds". When read aloud, the oral prompt will emphasize the phrase "3 fewer tens than hundreds" by making it louder and slower compared to the rest of the sentence. Additionally, the text of the oral prompt to be displayed will be adjusted to emphasize "3 fewer tens than hundreds" by displaying it in a different mode (e.g., bold, color, highlight, underline, italics, etc.) compared to the rest of the sentence. In another example, the phrase "the tens digit is 3 less than the hundreds digit" may be read progressively slower, with more variation in intonation, allowing learners to infer from visual and / or audio cues that the difference between the hundreds digit and the tens digit is 3, and that the hundreds digit is 7.
[0087] Similarly, if the learner selects option D (i.e., the answer is 2636), the generated differentiated correction prompt will be "The digit has a 7 in the hundreds place," and the generated speech emphasis mask will be "Has a 7 in the hundreds place." When the oral prompt is read aloud, it emphasizes the phrase "has a 7 in the hundreds place" by making it louder and slower compared to the rest of the sentence. Furthermore, the text of the oral prompt to be displayed is adjusted to emphasize "has a 7 in the hundreds place" by displaying it in a different mode compared to the rest of the sentence (e.g., bold, color, highlight, underline, italics, etc.). In another example, the phrase "has a 7 in the hundreds place" may be read progressively slower and adjusted to apply prosodic stress to "hundreds place," so that the learner can infer from the visual and / or audio cues that the digit in the hundreds place is 7.
[0088] If the learner selects option E (i.e., the answer is 746), the differentiated corrected prompt generated is "The number has four digits," and the generated speech emphasis mask is "has four digits." The oral prompt, when read aloud, emphasizes the words "has four digits" by being louder and slower compared to the rest of the sentence. Furthermore, the text of the oral prompt to be displayed is adjusted to emphasize "has four digits" by being displayed in a different mode (e.g., bold, color, highlight, underline, italics, etc.) compared to the rest of the sentence. In another example, the words "has four digits" may be read progressively slower and adjusted to apply prosodic stress to "4," so that the learner can infer from the visual and / or audio cues to understand that the answer should consist of four digits (i.e., four digits).
[0089] When the learner selects the correct answer (in this case, option A), the generated oral prompt and speech emphasis mask become "Well done." This oral prompt may be at a higher pitch when read aloud to motivate the learner. In another example, there may be applause after the words "Well done" are read aloud. Instead of showing the learner the words "Well done!", an image or video (e.g., a GIF) could also be shown to motivate the learner. Modifying the prosody of the oral prompt for the correct answer is optional.
[0090] In another language learning example, the learner device is configured to present language pronunciation guides and / or highlight difficult sounds, words, and phrases for specific learners. For example, if a learner is assigned a task to read a clause and the clause contains a word that the learner has misread (from a previous task), the learner device may be configured to display a list of the problematic word and related audio samples for the learner to consult before beginning an attempt at the learning task. In yet another example, difficult words (determined from the learner's history) may be displayed in a different mode (e.g., bold, highlighted, underlined, colored, etc.) from the rest of the clause so that the learner can better recognize their pitfalls.
[0091] In summary, the examples of AI-assisted learning methods disclosed in this disclosure generally involve monitoring a learner's feedback (or responses) to a specific learning task and providing relevant help regarding the learner's difficulties with the said learning task. Helpful verbal instructions are provided to effectively and clearly convey accurate information to the learner and to help the listener stay focused on the task at hand. The prompt language generation model is configured to generate personalized help information according to the distribution of the students' learning abilities. If an error occurs, or before the learner attempts the learning task, a sophisticated synthesized speech is generated to assist the learner. The synthesized speech is created by modifying at least some of the prosodic parameters of the speech (e.g., intonation, temporal rhythm, and volume) to provide the learner with acoustic cues, so that the learner can easily identify hidden or implied meanings between texts (of the learning task) and be guided to correct their errors or to give the correct answer. Useful speech signals (or acoustic cues) can be achieved by using a speech recognition model (e.g., an ASR acoustic model) that helps achieve temporal alignment of differentiated oral prompts (including speech emphasis) at the syllable, letter, word, or phrase level. This speech recognition model may be a machine learning-trained model. Information regarding temporal alignment, a modified speech signal including speech emphasis from a text-to-speech (TTS) module, and a differentiated emphasis mask (speech emphasis) are input to a digital signal processor (DSP). The DSP processes these inputs to generate a digitized, playable speech signal for the speech signal from the TTS module, which is temporally aligned according to the pitch, tempo, and / or volume information specified in the speech emphasis.
[0092] Experimental data: Table 2 below shows a comparison of the examples provided in this disclosure with systems operating using human tutors and existing technologies. [Table 5]
[0093] "Rel / Inc" represents the "relative increase" in the phoneme error rate. Table 2 illustrates the limitations of a conventional AI tutor (LJSpeech). In the smart AI tutor according to an embodiment of the present invention, the quality of oral feedback in differentiated oral prompts comprises at least two aspects: expressiveness and intelligibility. Existing TTS modules (such as LJSpeech) always speak in an unexpressive, monotonous voice, from which learners cannot perceive implicit meaning. The loss of intelligibility is significant compared to the voice of a human tutor. The loss of intelligibility is assessed by the relative increase in the phoneme error rate (speech recognition). On the other hand, the AI tutors of the embodiments of this disclosure have a low trade-off between expressiveness and intelligibility (relatively <5% compared to existing TTS modules) and can provide implicit meaning in oral prompts (like a human tutor).
[0094] Examples of this disclosure may have the following characteristics: Reference numerals in parentheses refer to reference numerals of elements in the figures.
[0095] A method and apparatus for generating one or more differentiated oral prompts for learning, the method or apparatus comprising: presenting a learning task (e.g., 102 in Figure 1, 202 in Figure 2) to a user (or learner) on a display; receiving user feedback on the learning task (e.g., 104 in Figure 1, 204 in Figure 2; e.g., user's answer); determining a diagnostic input (e.g., 228 in Figure 2) based on the user feedback; a differentiated oral prompt (e.g., 575 in Figure 5; e.g., a text string); and a differentiated speech enhancement mask associated with the differentiated oral prompt (e.g., 238 in Figure 2; e.g., indicating pitch, rhythm, volume and / or tempo settings). The performance involves performing a diagnosis on a diagnostic input using a prompt generation model (e.g., 115 in Figure 1, 230 in Figure 2, 515 in Figure 5) to generate an emphasis mask vector that may contain data, inputting the differentiated oral prompt and the differentiated speech emphasis mask into a text-to-speech module with prosodic control (e.g., 125 in Figure 1, 625 in Figure 6) to generate an audio signal for reading a differentiated oral prompt with the specified speech emphasis in the differentiated speech emphasis mask, and converting the audio signal into a format for reading a differentiated oral prompt with the specified speech emphasis via an audio device.
[0096] A method or apparatus for generating one or more differentiated oral prompts for a user may further comprise performing prosodic control at the word level, character level, or phrase level of the differentiated oral prompts.
[0097] Regarding a method or apparatus for generating one or more differentiated oral prompts to a user, the method or apparatus may further comprise using a speech recognition model (e.g., 535 in Figure 5, e.g., an ASR acoustic model) to achieve temporal alignment between differentiated oral prompts and speech emphasis at the word level, character level, or phrase level, the speech recognition model may be a model trained by machine learning, and information regarding temporal alignment and speech emphasis may be input to a digital signal processor (DSP) (e.g., 130 in Figure 1) to generate an audio signal according to pitch, tempo, and / or volume information specified in the speech emphasis. The DSP may also take an audio signal generated by a text-to-speech module as input.
[0098] The prompt generation model used to perform a diagnosis on a diagnostic input and generate differentiated verbal prompts may be a machine learning-trained model, and each differentiated verbal prompt may be specific to the learning task presented to the learner.
[0099] The speech recognition model may also be configured to use a Short-Time Fourier Transform (STFT) to segment the speech signal from the text-to-speech module into a first set of frames (e.g., 632 in Figure 6), where each frame corresponds to the spectral characteristics of an individual utterance (e.g., a word, phrase, character, syllable, etc.).
[0100] A first set of frames may be associated with information about temporal alignment, during temporal alignment, at least some of the prosodic parameters of the first set of frames are modified based on the speech enhancements specified in the differentiated speech enhancement mask so as to form a second set of frames (e.g., 640 in Figure 6).
[0101] A second set of frames may be processed by an inverse short-time Fourier transform (ISTFT) to generate an audio signal that can be converted into a format for reading through an audio device (e.g., an audio speaker).
[0102] With respect to a method or apparatus for generating one or more differentiated oral prompts for a user, the method may include checking a diagnostic input to determine whether a differentiated oral prompt is required for a learning task, before commencing the step of inputting a differentiated oral prompt and a differentiated speech emphasis mask into a text-to-speech module having prosodic control to generate an audio signal for reading a differentiated oral prompt having the speech emphasis specified in the differentiated speech emphasis mask.
[0103] The method may include checking a diagnostic input before starting step (e) to determine whether differentiated verbal prompts are required for the learning task.
[0104] The method may include presenting the user with visual cues on a display to assist in completing or correcting a learning task. The visual cues may be displayed in text format, and portions of the displayed text may be emphasized according to a differentiated speech emphasis mask, and portions of the text may be displayed in uppercase letters, bold, underlined, italicized, and / or in a different color from the rest of the text.
[0105] Diagnostic input may be determined based on the evaluation of speech (for example, for language learning) in user feedback.
[0106] Diagnostic inputs may be determined based on results obtained from error source detection in user feedback (e.g., for mathematical subject learning).
[0107] The diagnostic input may be determined based on the verification of one or more facts provided in user feedback (e.g., for scientific subject learning).
[0108] The specified speech emphasis may be an audio cue to assist the user in completing or correcting a learning task.
[0109] A differentiated voice emphasis mask may be configured according to one or more emotion settings, including: Instructional and / or encouraging modes for specifying vocal emphasis according to i-tension; Delightful mode and / or Joyful mode for specifying vocal emphasis according to the stress-stress rhythm; A condemnation mode for specifying vocal emphasis according to the rhythm of stress; An absolute mode for specifying speech emphasis according to a stress-weak-weak rhythm.
[0110] An apparatus for generating one or more differentiated oral prompts for learning, wherein the apparatus is (a) Presenting learning tasks to the user on a display, (b) Receiving user feedback on learning tasks, (c) Determining diagnostic inputs based on user feedback, Performing a diagnosis on a diagnostic input using a prompt generation model to generate differentiated oral prompts and differentiated speech enhancement masks associated with those differentiated oral prompts, wherein the differentiated speech enhancement masks include data indicating prosodic control. The differentiated oral prompt and the differentiated oral emphasis mask are input to a text-to-speech module with prosodic control in order to generate an audio signal for reading a differentiated oral prompt having the specified vocal emphasis in the differentiated vocal emphasis mask. The device comprises a processor for executing instructions in memory to control the device, which converts an audio signal into a format for reading aloud differentiated oral prompts with specified speech emphasis via an audio device.
[0111] In this disclosure, unless the context clearly indicates otherwise, the term “comprising” has the non-exclusive meaning of the word, “including at least,” rather than the exclusive meaning of “consisting only of.” The same applies to the corresponding grammatical changes for other forms of words such as "comprise" and "comprises."
[0112] While the present invention has been described in this disclosure in relation to several examples, embodiments, and implementations, the invention is not limited in that respect and includes a variety of obvious modifications and equivalent configurations that fall within the scope of the appended claims. Although the features of the invention are expressed in specific combinations within the claims, these features are intended to be arranged in any combination and order.
Claims
1. A method for generating one or more differentiated oral prompts for learning, (a) Presenting learning tasks to the user on a display, (b) Receiving user feedback on the learning task, (c) Determining the diagnostic input based on the user feedback, (d) Performing a diagnosis on the diagnostic input using a prompt generation model to generate the differentiated oral prompts and differentiated speech enhancement masks associated with the differentiated oral prompts, wherein the differentiated speech enhancement masks include data indicating prosodic control, (e) Inputting the differentiated oral prompt and the differentiated oral prompt mask into a text-to-speech module having prosodic control to generate an audio signal for reading aloud the differentiated oral prompt having the vocal emphasis specified in the differentiated vocal emphasis mask, (f) A method comprising converting the audio signal into a format for reading aloud the differentiated oral prompt having the specified voice emphasis via an audio device.
2. The method according to claim 1, further comprising performing prosodic control at the word level, character level, or phrase level of the differentiated oral prompts.
3. The aforementioned method, The method according to claim 2, further comprising using a speech recognition model to achieve the temporal alignment of the differentiated oral prompts and speech emphasis at the word level, character level, or phrase level, wherein the speech recognition model is a model trained by machine learning.
4. The method according to claim 3, wherein the temporal alignment and the information relating to the speech enhancement are input to a digital signal processor (DSP) to generate the speech signal according to the pitch, tempo, and / or volume information specified in the speech enhancement.
5. The method according to claim 3 or 4, wherein the speech recognition model is configured to use a short-time Fourier transform (STFT) to segment the speech signal from the text-to-speech module into a first series of frames, each frame corresponding to the spectral characteristics of an individual utterance.
6. The method according to any one of claims 3 to 5, wherein the speech recognition model is used, a first set of frames is associated with information about temporal alignment, and during the temporal alignment, prosodic parameters of at least some of the first set of frames are modified based on speech enhancements specified in the differentiated speech enhancement mask so as to form a second set of frames.
7. The method according to claim 6, wherein the second series of frames is processed by an inverse short-time Fourier transform (ISTFT) to generate the audio signal to be converted to the format for reading through the audio device.
8. The method according to any one of claims 1 to 7, wherein the prompt generation model is a model trained by machine learning, and each differentiated verbal prompt is specific to the learning task presented to the user.
9. The method according to any one of claims 1 to 8, further comprising checking the diagnostic input to determine whether a differentiated verbal prompt is required for the learning task before commencing step (e).
10. The method according to any one of claims 1 to 9, further comprising presenting the user with visual cues on the display to assist in the completion or modification of the learning task.
11. The method according to claim 10, wherein the visual cue is displayed in text format, a portion of the displayed text is emphasized according to the differentiated speech emphasis mask, and a portion of the text is displayed in uppercase letters, bold, underlined, italics, and / or in a different color from the rest of the text.
12. The method according to any one of claims 1 to 11, wherein the diagnostic input is determined based on the evaluation of the voice in the user feedback.
13. The method according to any one of claims 1 to 12, wherein the diagnostic input is determined based on the results obtained from the error source detection in the user feedback.
14. The method according to any one of claims 1 to 13, wherein the diagnostic input is determined based on verification of one or more facts provided in the user feedback.
15. The method according to any one of claims 1 to 14, wherein the specified speech enhancement is an audio cue for assisting the user in completing or correcting the learning task.
16. The distinguished voice enhancement mask is configured according to one or more emotion settings, Instructional and / or encouragement modes for specifying vocal emphasis according to the rhythm of the voice, Delightful mode and / or Joyful mode for specifying vocal emphasis according to stress-stress rhythm, A condemnation mode for specifying speech emphasis according to the rhythm of stress, The method according to any one of claims 1 to 15, further comprising an absolute mode for specifying speech emphasis according to a stress-stress-stress rhythm.
17. A device for generating one or more differentiated oral prompts for learning, wherein the device is The learning task is presented to the user on the display, Upon receiving user feedback on the aforementioned learning task, Based on the user feedback mentioned above, the diagnostic input is determined. A prompt generation model is used to perform a diagnosis on the diagnostic input to generate the differentiated oral prompts and the differentiated speech enhancement masks associated with the differentiated oral prompts, wherein the differentiated speech enhancement masks include data indicating prosodic control. The differentiated oral prompt and the differentiated oral emphasis mask are input to a text-to-speech module having prosodic control in order to generate an audio signal for reading aloud the differentiated oral prompt having the vocal emphasis specified in the differentiated vocal emphasis mask. A device comprising a processor for executing instructions in memory for controlling the device to convert the audio signal through an audio device into a format for reading aloud the differentiated oral prompt having the specified voice emphasis.